论文精选

AutoCompact:训练智能体自主决定上下文压缩时机

精选理由

一篇教智能体自己决定何时压缩上下文的论文,SWE-bench 上提了 9.2 个点,做 Agent 上下文管理的可以看看思路。

AutoCompact 训练智能体自己判断何时压缩上下文、保留哪些工作状态以及如何恢复任务。流程上先由 judge 审查基础智能体的压缩决策并修正,再用修正后的轨迹做 SFT,随后用任务成功奖励做 RL,把编码与压缩能力联合训练。该方法在 SWE-bench Verified 上将 pass rate 提高 9.2 点,在 SWE-PolyBench Verified 上提高 5.0 点。即使上下文窗口扩到 256K 且不溢出,收益依然存在,说明学习式压缩的价值不限于缓解空间不足。

原文 · elvis

First AutoHarness, then AutoContext, now AutoCompact. I am seeing a rising trend of work that trains models to natively support more of what the harness does. This work specifically trains agents to decide for themselves when to compact. Reminds me of the new paper from Meta that trains models to manage context natively. But how good is this approach? AutoCompact trains the agent to decide when to compact, what working state to keep, and how to resume. A judge first reviews the base agent's compaction decisions and replaces flawed ones before they execute. The corrected trajectories are used for SFT, then RL with task-success rewards trains coding and compaction together. Pass rates improve by 9.2 points on SWE-bench Verified and 5.0 points on SWE-PolyBench Verified. The gain holds even with a 256K window that never overflows, so learned compaction helps when context space is not the limit. It remains to be seen how this works at scale and how robust it is across harnesses. One interesting note from the authors is that this type of proactive compaction is a form of model-harness co-design: the harness provides the compaction mechanism, while the model learns when to invoke it, what to preserve, and how to continue afterward. Even more interesting is how to combine the rule-based compaction techniques already packaged in harnesses with more model-invoked proactive ones. Paper: arxiv.org/abs/2610.02163 Chat with Paper: academy.dair.ai/papers/autocom… 💬 19 🔄 5 ❤️ 46 👀 3432 📊 27 ⚡