论文

Microsoft 论文:无中心管理者的编程智能体团队表现更好

精选理由

Microsoft 发了篇论文,说编程智能体不用主控也能干活,128 个智能体自己认领任务、走 Git 合并,分数一路涨。

Microsoft 的论文提出 Agensh 框架,去掉主控智能体,让每个智能体自行认领子任务,在共享 Git 仓库中开发、测试、合并,并通过共享看板记录进展。在 ProgramBench 的 200 个任务中最难的 5 个上,团队规模从 1 个扩展到 128 个智能体,平均分数逐步提升。相比之下,主流多智能体编程工具都要经过 1 个主控智能体分发任务,可管理的辅助数量有限。论文所有实验使用同一模型,未报告大规模团队的成本。

原文 · rohanpaul_ai

Turns out coding agents don't need a boss:

New Microsoft paper finds that bigger teams of coding agents score higher and get there sooner when agents claim their own tasks without a central manager.

Even on the 5 hardest of ProgramBench's 200 tasks, scaling a manager-free team from 1 to 128 agents raised the average score at every step.

so for big jobs, add agents and let them coordinate through shared tools with no lead agent.

Popular multi-agent coding tools send all work through 1 lead agent, which can only manage so many helpers.

Agensh drops the lead agent. Each agent claims a sub-task, builds and tests it, merges it into a shared Git repo, and logs findings on a shared board.

Every run used 1 model, and the paper does not report what large teams cost.