NVIDIA 论文:终端智能体加大验证器投入比增加采样更有效
NVIDIA 新论文说终端命令别直接跑,先采样再验证,验证器要舍得花钱,Pass@1 直接从 50% 拉到 68%。
NVIDIA 发表关于终端智能体测试时计算的论文,核心做法是采样多条候选 shell 命令并在执行前验证。用 GPT-5.6 Sol 作为验证器从 8 个候选动作中挑选时,TerminalBench-Lite 的 Pass@1 从 50.0% 提升到 68.0%,而弱验证器下增加采样几乎没有收益。Mid-Harness 方法不改动生成器和 harness,在两者之间做验证。当 TMAX-9B 小模型自验证时,成对比较效果最好,再把强验证器蒸馏进去还能进一步提升。
Banger paper from NVIDIA on test-time compute for terminal agents.
The finding is that you should sample several candidate shell commands, verify them before running one, and spend more on the verifier than on extra samples.
With a GPT-5.6 Sol verifier choosing among 8 sampled actions, TerminalBench-Lite Pass@1 rises from 50.0% to 68.0%. With a weak verifier, extra samples add almost nothing.
Mid-Harness leaves the generator and harness unchanged and works between them. When a small TMAX-9B model verifies its own candidates, pairwise comparison works best, and distilling the strong verifier into it helps further.
Combining action sampling with trajectory sampling reaches higher success at lower estimated token cost than sampling full trajectories alone.
Paper: https://t.co/liDJ6AKnA0