论文多源确认精选

Nvidia 论文 VERA:模型训练与技能文件交替更新提升智能体长任务表现

精选理由

Nvidia 出的 VERA 方法挺实在:训练模型和改技能文件交替来,逐步打分定位哪一步出错,9B 智能体从 43.3 提到 69.1。

Nvidia 发布论文 VERA,针对长链路多步骤智能体任务提出交替更新方案:轮流训练模型和编辑其技能文件,并用真实证据的逐步评分决定每轮修复方向。VERA 构建了超过 9,000 个可重启沙箱,对工作流的每一步对照真实文件和日志打分,而不是只看最终结果。在医学研究基准上,9B 智能体采用两种更新方式后得分 69.1,仅用技能编辑为 56.1,仅用训练为 43.3。论文指出只训练模型或只改 harness 会损失约一半收益。

原文 · rohanpaul_ai

New Nvidia paper shows agents for long, multi-step work improve most when you alternate between training the model and editing its skill files, using step-by-step scores from real evidence to pick each fix.

Training only the model or only the harness leaves about half the gain on the table, compared with updating both in alternating rounds.

Most environments score only the final result, which hides which step broke. VERA builds over 9,000 restartable sandboxes that check each step against real files and logs, then improves the agent in rounds.

VERA turns benchmark runs into over 9,000 restartable sandboxes where each step of a long workflow gets its own checklist score from real evidence.

On a medical research benchmark, a 9B agent scored 69.1 with both kinds of updates, versus 56.1 with skill edits alone and 43.3 with training alone.

Score each step against real artifacts, and let those scores decide whether the next fix goes into the model or its skills.