微软发布 CorpusMap:让智能体跨文档检索更省 token
微软发了篇论文叫 CorpusMap,提前把文档集合里的实体建好索引页,智能体检索时准确率涨最多 11.7 分,token 还省一半多。
微软及合作者发表论文 CorpusMap,面向运行大规模文档集合检索的智能体。该方法提前解析集合中的重复实体,为每个实体建立页面并链接到所有提及它的文档,原始文档保持不变。智能体读取文档后可沿实体跳转到相关文档,避免重复搜索同一证据。在 7 个模型和三个基准上,答案质量提升 6.4 至 11.7 分,输入 token 减少 34% 至 57%,并超过 LLM Wiki 层和另外三种导航层。地图构建无需 LLM 调用,新文档到达时可增量更新。
Banger paper from Microsoft and colleagues.
If you run agents that search a large document collection, this one is worth your time.
(bookmark it)
They introduce CorpusMap, which resolves recurring entities across the collection in advance and gives each entity a page that links to every document that mentions it.
The original documents stay in place.
The agent reads a document, follows an entity to related documents, and avoids searching for the same evidence again.
Across 7 models and three benchmarks, answer quality goes up 6.4 to 11.7 points while input tokens drop 34% to 57%.
It also beats an LLM Wiki layer and three other navigation layers.
The map can be built without LLM calls and updated as new documents arrive.
Paper: https://t.co/GimGee8u8V
Chat with Paper: https://t.co/0GOcC169at