Meta 等机构新论文:分支化自我改进的 Agent Harness 优化
Meta 和 Duke 的论文把 agent harness 自动调优拆成专精分支再加路由,Olympiad 数学准确率从 46% 拉到 62%,做 agent 自动调优的可以看看。
Meta、Duke 和加州大学发布关于 self-improving agent harness 优化的论文。Harness 是包裹 LLM 的代码,控制工具、检索和自检环节。论文将 harness 搜索拆成 2 个分支,每个分支保留自己更擅长的题目并记录各自的优化笔记,再用 router 把任务分发给最合适的 harness。在 Olympiad 数学上使用 Gemini 3 Flash,准确率从 Meta-Harness 的 46.0% 提升到 62.0%,并在全部 4 个测试设定中超过 Meta-Harness。
New paper from Meta, Duke, California Univ on self-Improving Agent's Harness Optimization
When an AI tunes your agent's harness, split the search into specialized branches and route each task to the best fit, which beat Meta-Harness in all 4 test settings.
Giving each tuning branch its own problems and its own notes on what worked produced harnesses with different strengths, and a router turned those strengths into higher scores.
A harness is the code around an LLM that controls its tools, retrieval and self-checks. Meta-Harness has an AI rewrite it in a loop, but every version is scored on the same problems, so search sticks to 1 path.
Each of 2 branches keeps the practice problems it solves better than the other and writes its own notes on what worked. In math, 1 branch learned to verify answers, while the other learned to build full derivations.
With Gemini 3 Flash on Olympiad math, accuracy rose from 46.0% with Meta-Harness to 62.0%.
If you auto-tune agents, keep several specialized harnesses and route between them rather than betting on 1 winner.