论文

Harness-Zero:把专用 Agent Harness 的能力蒸馏进模型权重

Harness-Zero: Harness Distillation via Agent-as-Harness

精选理由

把专用 harness 的本事直接练进模型权重,部署时甩掉笨重框架,成绩反而从 23.3% 涨到 44.3%。

arXiv 论文提出 Harness-Zero,用 agent-as-harness 方式做 harness 蒸馏:一个 harnessing agent 在目标 harness 的动作空间内先修正学生模型的回答,再把专用 harness 的引导转化为训练示范。在知识工作、工具使用和科学三个领域的实验中,部署时移除专用 harness 后,基础模型的宏观平均任务成功率从 23.3% 升到 44.3%,还超过了仍挂着该 harness 时的 41.7%。对前沿 LLM 使用同一个 evolved harness 时,agent-as-harness 的效果优于 code-as-harness。在三个领域共 28 种行为模式上,该方法平均恢复了 82.3% 由 harness 引入而基础模型缺失的行为。

原文 · arXiv cs.AI

Harness-Zero: Harness Distillation via Agent-as-Harness

Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a general-purpose agent must either settle for a suboptimal shared harness or route among an ever-growing set of specialized ones. We therefore study agent harness distillation: using a domain- or instance-optimized harness as training-time guidance and transferring the behaviors it induces into model weights, so that its gains survive under a single fixed target harness. The challenge is that the two harnesses differ in action space and available information, so guidance from the optimized harness cannot serve directly as supervision for the target one. We introduce Harness-Zero, which enables harness distillation through agent-as-harness. Guided by the optimized harness, a harnessing agent corrects student responses before execution in the target harness's action space, turning harness guidance into training demonstrations. Fine-tuning on the resulting trajectories internalizes harness-induced behavior into the model, so the specialized harness can be removed at deployment. Our experiments spanning knowledge work, tool use, and science domains show that: (1) For frontier LLMs using the same evolved harness, agent-as-harness outperforms code-as-harness. (2) With the specialized harness removed at deployment, Harness-Zero improves the base model's macro-average task success from 23.3% to 44.3%, even exceeding the 41.7% it reaches with that harness still attached. (3) Harness-Zero recovers harness-induced behaviors absent from the base model, with 82.3% average recovery across 28 patterns in the three domains.