Looped Mixture of Experts扩展定律研究
Scaling Laws for Looped Mixture of Experts
这篇论文提出了首个联合建模循环和稀疏性的扩展定律,万亿token规模实验显示其能实现2倍参数效率提升。
研究人员提出Loop Scaling Laws,首次联合建模循环和稀疏性,结合模型规模和数据。该定律通过有界稀疏条件循环映射,量化了循环带来的有效参数增益及稀疏性如何提升此增益。在保留损失预测上,该定律比先前方法更准确,并能恢复标准密集和MoE扩展定律作为特例。下游评估显示稀疏性提供约3倍活跃参数效率,循环推理带来约2倍总参数效率,联合扩展进一步推进性能边界。
Scaling Laws for Looped Mixture of Experts
Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. The laws predict the held-out loss of looped models more accurately than prior alternatives, and recover the standard dense and MoE scaling laws as special cases. Beyond prediction, the fitted laws provide a principled foundation for designing looped MoE models under compute and memory constraints. Downstream evaluations further demonstrate the complementary benefits of the two axes: sparsity delivers ~3x active-parameter efficiency, recurrence yields ~2x total-parameter efficiency on reasoning, and joint scaling further advances the performance frontier. As a practical extension, we show these gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence.