论文

SourceLearn:开发特定源学习能力

From Knowledge Access to Source Learning: Developing Source-Specific Competence

精选理由

研究人员提出SourceLearn,让AI能从反复使用同一源中学习,而非简单重复访问,在多个基准测试中大幅提升性能。

SourceLearn是一种新方法,通过自定向源学习和任务引导源学习两种机制,构建并逐步完善持久源模型。该方法在五个基准测试和三种LLM后端上表现优异,在15种设置中的13种中取得最佳性能,比混合RAG提升最高达22.6分。源模型能够捕获对权威源的可重用理解,包括其知识结构、解释和应用方式。

原文 · arXiv cs.LG

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-specific competence over a persistent authoritative source. We represent this competence with a persistent source model that captures reusable understanding of the source, including how its knowledge is structured, interpreted, and applied. To construct and progressively refine such models, we propose SourceLearn, which combines two complementary learning mechanisms. Self-Directed Source Learning identifies what remains incompletely understood and adaptively revisits the source, while Task-Guided Source Learning uses downstream experience to reveal local representational gaps and recurring needs in how source knowledge should be organized. In both cases, learning signals determine what should be reconsidered, while persistent updates are reconstructed from the authoritative source. Across five benchmarks and three LLM backends, SourceLearn achieves the best performance in 13 of 15 settings, with gains of up to 22.6 points over Hybrid RAG and substantial overall improvements over static source representations and experience-based memory baselines.