论文精选73°

Meta发布智能体长程运行控制研究

精选理由

Meta研究如何让AI智能体更高效地长程运行,控制器让GPT-5.5在相同资源下表现大幅提升。

Meta Superintelligence Labs发布论文,提出智能体长程运行控制方法。在ProgramBench基准测试中,GPT-5.5配合专用控制器准确率从63.7%提升至71.5%,超越Codex的58.0%。控制器通过维护运行摘要、预估工作价值,在ProofBench、ARC-AGI-2和LongCoT-mini基准上平均提升3.6至4.2分。

原文 · DAIR.AI

Super interesting paper from Meta Superintelligence Labs on controlling long agent runs.

They use the same workers and same budget, and ProgramBench goes from 63.7% to 71.5% with GPT-5.5 when a dedicated controller decides what work to run next.

Codex scores 58.0%.

Here is how it works:

The controller keeps a short summary of the run and leaves full worker outputs in memory. Each cycle it updates that summary, proposes next steps, estimates what each one is worth under the remaining budget, and sends the chosen work to workers with the earlier outputs they need.

On ProofBench, ARC-AGI-2 and LongCoT-mini it adds 3.6 to 4.2 points over direct control, averaged over three frontier models.

Paper: https://t.co/34PuaVLnkr