论文精选

Microsoft 发布 ActiveSaddler:按未修复失败模式自动优化 agent harness

精选理由

微软这篇论文讲怎么优化 agent 的 harness,核心是把预算花在还没修好的失败上。GAIA2 提了 4.4 个点,成本还降到四分之一,做 agent 自动调优的可以看看。

Microsoft 发表一篇关于 agent harness 自动优化的论文,提出 ActiveSaddler 方法。该方法不再使用固定任务列表,而是追踪失败模式,优先处理最值得修复的失败,或探索未见过的任务以发现新失败模式。在 GAIA2 基准上,ActiveSaddler 将测试通过率提升 4.4 个百分点,在 Terminal-Bench 2.0 上提升 7.5 个百分点。达到 58.5% 的 GAIA2 dev 准确率只花费 298 美元,而固定任务顺序需要 1360 美元。

原文 · rohanpaul_ai

New Microsoft paper on Automated harness optimization for agents.

Most harness auto-tuners focus on how to patch prompts and tools, but which tasks produce the feedback also changes how good the final harness gets.

But you will get stronger AI agents when you pick training tasks based on which failures are still unfixed, so stop feeding them a fixed task list.

ActiveSaddler tracks failure patterns and works on the one most worth fixing, or tries unseen tasks to find new ones.

On the same optimizer, ActiveSaddler raised test pass rates by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0. Reaching 58.5% GAIA2 dev accuracy cost $298, versus $1,360 with a fixed order.

If you auto-tune an agent, aim your run budget at the failures that are still open.