论文精选

UT Austin 研究:压缩上下文可能让智能体更慢而非更快

精选理由

UT Austin 跑了3.5万次实验发现:为了省 token 压缩上下文,可能反而让智能体慢 20%-80%,而且策略换模型就失效。

UT Austin 团队在 SWE-bench Verified 和 Terminal-Bench 上跑了近 35000 次编码智能体实验,分别改变压缩方式、触发时机和删除比例三个变量。在 Terminal-Bench 上用 Qwen 测试时,压缩到约三分之一 token 的策略比保留完整上下文慢 20% 到 80%。按步触发的压缩每步省 token 最多,但模型调用次数多出 10% 到 27%;阈值触发策略能削减 22% 到 55% 的 token,调用次数接近完整上下文。同一策略在 Qwen 上表现良好,却让 Devstral 降到 38.7% 且变慢,选压缩策略前需要针对具体模型测延迟和成本。

原文 · elvis

Great overview of context compression in LLM Agents And one interesting, unexpected finding. If your compaction policy is tuned to cut tokens, it may be making your agent run slower. Great study from UT Austin on context compression in coding agents. They ran nearly 35,000 agent runs on SWE-bench Verified and Terminal-Bench and varied three decisions separately. These are how context is compressed, when compression triggers, and how much is removed. On Terminal-Bench with Qwen, policies that use about a third of the tokens can take 20% to 80% longer than keeping full context. Step-triggered policies cut the most tokens per step but need 10% to 27% more model calls. Threshold-triggered policies cut tokens by 22% to 55% with call counts close to full context. Results also differ by model. A policy that works well for Qwen drops Devstral to 38.7% and makes it slower, so measure latency and cost per model before choosing one. Paper: arxiv.org/abs/2609.32961 Chat with Paper: academy.dair.ai/papers/beyond-… 💬 10 🔄 1 ❤️ 28 👀 1561 📊 13 ⚡