Stanford 论文 DeLM:去中心化多智能体在编程任务上跑赢 Claude Code
Stanford 把多智能体编程的中央调度改成了共享任务队列,在 Codex 和 Claude Code 上实测快 2.49 倍,还开源了插件,可以直接试。
Stanford 论文提出 DeLM,用共享上下文和任务队列替代中央协调智能体。智能体独立领取任务,实时共享发现,避免重复劳动和相互等待。在 Terminal-Bench 4.0、DeepSWE v1.1 和 ProgramBench 的长程任务上,DeLM 比 vanilla Claude Code 和 Codex 基线最快快 2.49 倍,准确率最多高出 19.2 个百分点。同样 120 分钟预算内,ProgramBench 测试通过率最多提升 19.9 个点。代码和 720 条轨迹已开源,提供可直接在 Codex 和 Claude Code 中使用的插件。
A new Stanford paper just proved that decentralized multi-agent systems can beat Claude Code and Codex on complex coding tasks, running up to 2.49× faster!
DeLM tackles a key bottleneck in multi-agent systems: time wasted repeating work and waiting on other agents.
DeLM replaces the central coordinating agent with a shared context and task queue. Agents pick up tasks independently, share discoveries as they happen, and build on each other’s progress. When one agent finds a solution or hits a dead end, the others can use that information immediately.
DeLM beats state-of-the-art in both speed and accuracy on long-horizon tasks from Terminal-Bench 4.0, DeepSWE v1.1, and ProgramBench:
- Up to 2.49× faster execution than the vanilla Claude Code and Codex baselines - Up to +19.2 percentage points in accuracy over the vanilla Claude Code and Codex baselines - Up to +19.9 points in ProgramBench test pass rate within the same 120-minute budget
Built on Codex and Claude Code. Code and 720 trajectories are available so you can explore how the agents collaborate! An open-source plugin lets you try DeLM directly in Codex and Claude Code.