技巧精选

Milvus 实测:Jev 能否替代 RAG 重排器

𝗖𝗮𝗻 𝗝𝗲𝘃 𝗿𝗲𝗽𝗹𝗮𝗰𝗲 𝗮 𝗿𝗲𝗿𝗮𝗻𝗸𝗲𝗿 𝗶𝗻 𝗥𝗔𝗚? We tested three setups on the same Mi...

精选理由

Jev 重排效果超过 qwen3.7,但延迟高 10 倍、成本近 7 倍,Milvus 建议只用于离线场景。

Milvus 团队在 80 条 SciFact 查询上对比三种方案:不重排、qwen3.7-text-rerank 和 Jev。qwen3.7-text-rerank 将 nDCG@10 提升 0.0446,Jev 提升 0.0778,排序质量三者最佳。但 Jev 的 P50 重排延迟是 qwen 的 10.2 倍,单次运行成本估算约为其 6.7 倍,原因是 30 个候选需要 30 个并发请求。用 P(Yes) 概率做 0.5 阈值过滤时,top-5 精度为 49.4%,同时 18.8% 的黄金相关文档被误滤掉。结论是 Jev 适合离线检索、数据清洗和评估,难以直接用于延迟敏感的在线 RAG。

原文 · Milvus

𝗖𝗮𝗻 𝗝𝗲𝘃 𝗿𝗲𝗽𝗹𝗮𝗰𝗲 𝗮 𝗿𝗲𝗿𝗮𝗻𝗸𝗲𝗿 𝗶𝗻 𝗥𝗔𝗚? We tested three setups on the same Mi...

𝗖𝗮𝗻 𝗝𝗲𝘃 𝗿𝗲𝗽𝗹𝗮𝗰𝗲 𝗮 𝗿𝗲𝗿𝗮𝗻𝗸𝗲𝗿 𝗶𝗻 𝗥𝗔𝗚? We tested three setups on the same Milvus shortlist: 𝗻𝗼 𝗿𝗲𝗿𝗮𝗻𝗸𝗶𝗻𝗴, 𝗾𝘄𝗲𝗻𝟯.𝟳-𝘁𝗲𝘅𝘁-𝗿𝗲𝗿𝗮𝗻𝗸, 𝗮𝗻𝗱 𝗝𝗲𝘃. On 80 SciFact queries, qwen3.7-text-rerank improved nDCG @10 by 𝟬.𝟬𝟰𝟰𝟲 over no reranking, while Jev improved it by 𝟬.𝟬𝟳𝟳𝟴. Jev ranked best of the three, but its P50 reranking latency was 𝟭𝟬.𝟮× 𝗵𝗶𝗴𝗵𝗲𝗿 𝘁𝗵𝗮𝗻 𝗾𝘄𝗲𝗻, with an estimated cost per run of 𝟲.𝟳×. Why does Jev behave differently? Its interface is built around 𝘀𝘁𝗮𝘁𝗲 + 𝗾𝘂𝗲𝘀𝘁𝗶𝗼𝗻𝘀. For reranking, we keep the relevance question fixed and change the state for each query-candidate pair: “𝗜𝘀 𝘁𝗵𝗶𝘀 𝗰𝗮𝗻𝗱𝗶𝗱𝗮𝘁𝗲 𝗿𝗲𝗹𝗲𝘃𝗮𝗻𝘁 𝘁𝗼 𝘁𝗵𝗲 𝗾𝘂𝗲𝗿𝘆?” Jev returns a structured Yes/No judgment with probabilities, and we use 𝗣(𝗬𝗲𝘀) as the relevance score for sorting. In this experiment, 30 candidates meant 30 independent Jev requests running concurrently, so the latency includes that API overhead. We also tested using probability directly for filtering. At a 𝟬.𝟱 𝘁𝗵𝗿𝗲𝘀𝗵𝗼𝗹𝗱, top-5 precision reached 𝟰𝟵.𝟰%, but 𝟭𝟴.𝟴% 𝗼𝗳 𝗴𝗼𝗹𝗱-𝗿𝗲𝗹𝗲𝘃𝗮𝗻𝘁 𝗱𝗼𝗰𝘂𝗺𝗲𝗻𝘁𝘀 𝘄𝗲𝗿𝗲 𝗳𝗶𝗹𝘁𝗲𝗿𝗲𝗱 𝗼𝘂𝘁. That makes Jev attractive for offline retrieval, data cleaning, and evaluation, but much harder to use as-is for latency-sensitive online RAG. Next, we’re testing whether batching multiple candidate questions into one Jev request can bring that latency down. 💬 0 🔄 0 ❤️ 1 👀 92 📊 1 ⚡