技巧73°

NVIDIA研究提升AI代理可靠性方法

精选理由

NVIDIA让AI代理多选一判断,不重训练就提18%成功率,终端操作更可靠。

NVIDIA研究人员提出新方法,让AI代理在每个步骤生成8个选项,由判断模型选择最佳执行。使用前沿模型作为判断者,成功率从50%提升至68%,无需重新训练模型。小型模型自我判断时提升效果有限,判断能力不足时增加选项效果不大。

原文 · rohanpaul_ai

Nvidia's new paper tell before giving your agent more options, give it a better judge.

A single bad command can derail an AI agent, so have it draft a few options and let a smart judge pick before acting.

And in this way, You can make an AI agent far more reliable without retraining it.

AI agents that work in a terminal usually run the 1st command they come up with. A single bad move, like installing the wrong package, can throw off every step after it.

NVIDIA researchers had a small agent draft 8 options at each step and let a judge choose which to run. With a strong frontier model as the judge, its success rate jumped from 50% to 68%, no retraining needed.

When the small model judged its own drafts, the gains were much smaller. More options don't help much if the judge can't tell them apart.