论文精选73°

Meta研究显示带管理器的智能体随算力提升表现

精选理由

Meta发现给AI智能体加个'预算经理'比单纯堆算力更有效,编程任务成绩提升明显。

Meta发表新论文,提出元推理智能体架构。该架构由工作器执行任务,独立控制器决定资源分配。在12项对比测试中,带管理器的智能体全面胜出。使用GPT-5.5编程基准测试,三倍预算使分数从64.1%提升至71.5%,而普通智能体停滞在64%左右。

原文 · rohanpaul_ai

New Meta paper shows that long-running agents keep getting better with more compute when a separate manager decides how to spend it.

More compute gives an agent more choices: what to try next, what to trust, when to stop. Many agents make those calls on the fly, so extra budget can go to waste.

Their fix hands those calls to a separate manager that thinks them through, while workers do the actual task.

They built a Meta-Reasoning Agent: workers do the task, and a separate controller decides what comes next. It keeps a short progress summary, weighs options against the remaining budget, and picks which past results each worker sees.

At the largest budget, this setup beat an otherwise identical agent without the manager in all 12 head-to-head tests. With GPT-5.5 on a coding benchmark, tripling the budget lifted its score from 64.1% to 71.5%, while the other agent stalled near 64%.

For long-running agents, don't just add compute: spend some on a manager that decides where the rest goes.