Steiner任务分类框架揭示多智能体规模化的结构性差异
Multi-agent Scaling Across Disjunctive and Compensatory Tasks
一篇挺有意思的研究:多智能体堆人海到底有没有用,取决于任务类型,13 个模型、30 个智能体规模的实验说得很细。
arXiv 论文引入 Steiner 群体任务分类作为分析多智能体 LLM 规模化的框架,区分析取式与补偿式任务。实验覆盖 13 个开源权重模型、最多 30 个智能体的团队。在析取式任务上,至少一个智能体正确的概率随团队规模提升 5-20 个百分点,但直接多数投票几乎无法兑现这一潜力,多轮修订提升明显且一个同伴与 29 个同伴收益接近。在 Fermi 估计任务上,模型内共享的条目级偏差约占 87% 的平方误差,平均只能降低约 6% 的误差,混合模型家族有帮助但不超过最强成员。
Multi-agent Scaling Across Disjunctive and Compensatory Tasks
Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model's modal answer, and averaging converges to the model's item-level bias. Across selected representative benchmarks, 13 open-weight models, and teams of up to 30 agents, we find qualitatively different scaling behavior. On disjunctive tasks, the probability that at least one agent is correct grows by 5-20 points with team size, but plurality voting over agents that answer directly realises almost none of this potential, as the model predicts to within 0.5 points on average. Multi-round revision raises accuracy considerably, yet the gain is nearly the same with one peer as with 29. In contrast, scaling provides little benefit on Fermi estimation, despite its natural suitability for aggregation: item-level biases shared across the samples of a model account for about 87% of the squared error, so averaging reduces error by only about 6%. Combining model families helps on Fermi estimation but does not surpass the strongest member on disjunctive tasks. These results show that task structure, together with the mechanism combining member outputs, is a fundamental determinant of team scaling.