Foil 方法提出循环 MoE 训练新设计:扁平化专家并解绑注意力
How to Loop MoE: Flatten the Experts, Untie the Attention
一篇教你怎么把 MoE 和循环 Transformer 结合的论文,Foil 把专家层数减半、循环翻倍,同等算力下预训练损失更低,代码也开源了。
Looped MoE 将循环 Transformer 与稀疏 MoE 结合,通过重复使用同一组层来提升固定参数规模的模型表现。论文提出 Foil 方法:在专家参数和每 token 专家计算量不变的前提下,将专家层数减半、每层专家数翻倍、循环次数翻倍,使每次路由从更大的专家池中选择;同时为每次循环配备独立的注意力参数,专家与路由器保持共享。实验显示在 100B token 预训练时,最扁平化的 Foil 比未扁平化的循环基线低 0.012 nat,下游准确率持平或更好。消融实验还给出设计指导:循环与加宽专家层的收益相互放大,路由置信度比负载均衡更能反映专家使用是否健康。
How to Loop MoE: Flatten the Experts, Untie the Attention
Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil. With the expert parameters and the expert compute per token held fixed, Foil (1) flattens the experts, halving the expert layers, doubling the experts per layer and doubling the passes, so that every routing decision chooses from a larger pool, and (2) unties the attention, giving each pass its own attention parameters while the experts and routers stay shared. Experiments show that Foil clearly outperforms the unflattened looped baseline: at 20B tokens every Foil model has lower pretraining loss than the baseline; at 100B tokens the loss improves monotonically with the degree of flattening, the most flattened Foil ending 0.012 nat below the baseline at equal parameters and compute, with downstream accuracy on par or better; untying the attention also yields more balanced and more confident routing at equal shape. Our ablations analyse why Foil works and turn the findings into design guidance for looped MoE: the returns of looping and of widening the expert layers amplify each other, routing confidence tracks healthy expert use better than load balance, and a sparse looped MoE should therefore use more experts per layer and more passes. Code and configurations are available at https://github.com/SR-A-W/how-to-loop-moe.