技巧精选

Milvus 3.0 用 late interaction 做视觉 PDF 检索演示

精选理由

Milvus 3.0 的演示,教你检索带图表的 PDF 时别用单向量,用 late interaction 保住细节,适合看论文和报告的人。

Milvus 3.0 演示了针对视觉密集 PDF 的检索方案:查询生成 token embeddings 列表,PDF 页面切成视觉 patch embeddings,再用 MAX_SIM_COSINE 做 late interaction 排序,公式为 score(Q, P) = Σᵢ maxⱼ cosine(qᵢ, pⱼ)。每个查询 token 会先找到页面上最相关的区域,再汇总成页面级得分,避免图表、箭头和空间关系被压进单一向量而丢失。演示用的是 NASA Systems Engineering Handbook 的页面,在线 demo 已开放在 demos.milvus.io/embedding-list/。

图片来源 · Milvus
原文 · Milvus

Single-vector retrieval is a useful default. But on visually dense PDFs, it can become lossy: diagrams, labels, arrows, and spatial relationships get compressed into one representation, even when a query depends on one local detail. This Milvus 3.0 demo takes a different approach: Query → a list of token embeddings PDF page → a list of visual patch embeddings Ranking → late interaction with MAX_SIM_COSINE In one line: score(Q, P) = Σᵢ maxⱼ cosine(qᵢ, pⱼ) For each query token, Milvus finds the most relevant region on the page, then combines those best matches into a page-level score. This keeps fine-grained visual signals intact until the moment of comparison. The demo applies the approach to pages from the NASA Systems Engineering Handbook, making it easy to see visual PDF retrieval working end to end. If you work with engineering manuals, scientific papers, financial reports, or slide decks, it’s worth a look. Explore the live demo: demos.milvus.io/embedding-list/ Your browser does not support the video tag. 🔗 View on Twitter 💬 0 🔄 0 ❤️ 2 👀 113 📊 1 ⚡