Milvus 3.0现在可以直接索引外部数据源,无需移动原始数据,简化了向量数据库工作流。
Milvus 3.0新增External Collection功能,可直接索引Parquet、Apache Iceberg等外部数据源。该功能支持ANN向量搜索、标量和JSON过滤、BM25全文搜索及混合检索排序。对于批量生产的大型数据集,无需将源行复制到Milvus或维护同步管道。
𝗠𝗶𝗹𝘃𝘂𝘀 𝟯.𝟬 𝗰𝗮𝗻 𝗻𝗼𝘄 𝗶𝗻𝗱𝗲𝘅 𝗮𝗻𝗱 𝘀𝗲𝗮𝗿𝗰𝗵 𝗱𝗮𝘁𝗮 𝘄𝗶𝘁𝗵𝗼𝘂𝘁 𝗳𝗶𝗿𝘀𝘁 ...
𝗠𝗶𝗹𝘃𝘂𝘀 𝟯.𝟬 𝗰𝗮𝗻 𝗻𝗼𝘄 𝗶𝗻𝗱𝗲𝘅 𝗮𝗻𝗱 𝘀𝗲𝗮𝗿𝗰𝗵 𝗱𝗮𝘁𝗮 𝘄𝗶𝘁𝗵𝗼𝘂𝘁 𝗳𝗶𝗿𝘀𝘁 𝗺𝗼𝘃𝗶𝗻𝗴 𝘁𝗵𝗲 𝘀𝗼𝘂𝗿𝗰𝗲 𝗿𝗼𝘄𝘀 𝗶𝗻𝘁𝗼 𝗠𝗶𝗹𝘃𝘂𝘀. External Collection connects Milvus directly to data stored in Parquet, Apache Iceberg, Lance, Vortex, or a supported Milvus snapshot. You define the external source, map its columns to a Milvus schema, and create the indexes your application needs. Milvus then builds and serves the retrieval layer while the original records remain in the lake. From the application side, the experience stays familiar. Once the collection is refreshed and loaded, teams can use the same Milvus search and query APIs for: • ANN vector search • scalar and JSON filtering • BM25 and full-text search • hybrid retrieval and ranking For large, batch-produced datasets, this provides a production retrieval path without copying the source rows into Milvus or maintaining another synchronization pipeline. Keeping the lake as the shared data foundation is a key milvus.io/blog/milvus-3-… s 3.0’s Vector Lakebase architecture. Read more: https://t.co/VWDpFH7KLn 💬 0 🔄 0 ❤️ 0 👀 19 ⚡