论文提出按失败构成判断该改 harness 还是训练权重
Harness Evolution Hits a Ceiling: When Weight Training Should Begin
论文教你看 agent 失败原因来决定改运行时框架还是微调模型,Qwen3.5 上从 0.16 拉到 0.30,做 agent 的都该看看这套诊断方法。
一篇 arXiv 论文研究了长程 LLM agent 的两条改进路径:演化外部 harness 或训练模型权重。作者用首次触发的信号给失败轨迹打标签,区分过程失败(阻塞调用、循环、步数耗尽)与内容失败(交付结果质量差),前者由 harness 修复,后者留给权重训练。在 DeepPlanning 上,自演化 harness 将 Qwen3.5-4B 的 held-out 分数从 0.16 提到 0.30,Qwen3.5-9B 从 0.32 提到 0.44。用演化轨迹训练的 LoRA 适配器在原 harness 下为两个模型规模各加 0.13,4B 叠加后 held-out 分数翻倍以上,9B 上适配器单独即追平完整演化线,内容失败从四分之一降到二十分之一。该流程在 WebArena-Lite 的 117 个未见任务上额外提升 0.09。
Harness Evolution Hits a Ceiling: When Weight Training Should Begin
Improving a long-horizon LLM agent means evolving the harness around a frozen model or training its weights. We let a self-evolving harness make the system stronger first, then cross seed and evolved harnesses with base and trained weights to learn which gains the trained model keeps and which still need the runtime. We show that the right lever can be read off the agent's failure composition: labelling failed trajectories by the first signal that fires separates process failures (blocked calls, loops, exhausted step budgets) from content failures (a delivered plan that is poor). Harness evolution repairs the former, the behaviour it instils can be trained into the weights, and content failures are what weight training is for. On DeepPlanning, a self-evolving harness loop lifts the held-out score of Qwen3.5-4B from 0.16 to 0.30 and of Qwen3.5-9B from 0.32 to 0.44; for 4B, held-out delivery rises from 55% to 90% while content failures are left for the weights. LoRA adapters trained on evolved-harness trajectories internalise the gain: under the original harness they add +0.13 on held-out tasks for both sizes; on 4B they stack with the harness to more than double the held-out score, and on 9B the adapter alone matches the full evolution line, cutting content failures from a quarter of trajectories to one in twenty. A placebo adapter trained on answer-shuffled trajectories falls below the base model. The loop transfers to WebArena-Lite (+0.09 on 117 unseen tasks), where the gain lives in what the model sees and adapters do not add to it. The result is a diagnose-then-intervene rule applied twice: read the failure composition to choose between harness and weights, then read what the accepted edits changed to decide which gains to train in. Scores are four-rollout means against fresh anchors, same-night except where marked, across eight models from six families and two benchmarks.