论文精选

多智能体团队性能研究

More agents don't mean higher performance. There is a coordination bottleneck to consider. Not to m...

精选理由

AgentWorld论文揭示多智能体团队的协调瓶颈,Gemini 3 Flash表现最佳,但协调任务成功率仅12%

AgentWorld研究表明,多智能体团队中不到三分之一的行动有助于完成任务。研究将3至20个具有不同角色的LLM智能体放入游戏沙盒,执行50多轮任务。Gemini 3 Flash表现最佳,任务成功率达52.0%,其中协调任务成功率最低,仅为12%。

原文 · elvis

More agents don't mean higher performance. There is a coordination bottleneck to consider. Not to m...

More agents don't mean higher performance. There is a coordination bottleneck to consider. Not to mention the unnecessary costs. So how many of a multi-agent team's actions actually help it finish the task? In this AgentWorld paper, fewer than a third. AgentWorld puts 3 to 20 LLM agents with different roles into a game sandbox for tasks that run 50+ rounds. Agents can't see each other's internal state, so they have to coordinate through messages and shared plans. Gemini 3 Flash has the highest task success at 52.0%. Coordination tasks are the hardest category, at 12% success, and common failures include communication breakdowns, role confusion, and lost shared plans. Paper: arxiv.org/abs/2609.31590 Chat with Paper: academy.dair.ai/papers/agentwo… 💬 0 🔄 0 ❤️ 0 ⚡