论文

Elastic Expert Routing:软化 MoE top-k 路由边界,提升 OLMoE 与 Qwen3 微调效果

Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts

精选理由

一篇 MoE 训练论文:把固定 top-k 换成在 k 附近随机采样,Qwen3-30B-A3B 微调能多拿 2 分,推理成本不变。

论文提出 Elastic Expert Routing,将传统 MoE 训练中固定 top-k 专家选择改为在以 k 为中心的局部离散分布上随机采样活跃专家数。该方法把生硬的阈值变成渐进概率分布,期望计算成本与确定性训练相同,推理预算不变。在 OLMoE-1B-7B 和 Qwen3-30B-A3B 的监督微调中,下游 macro 平均分别提升 0.84 和 2.02 分;从头预训练场景平均超过静态 top-k 基线 1.6 分。

原文 · arXiv cs.LG

Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts

Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token. However, traditional training paradigms enforce a static choice of top-$k$ experts, which converts a continuous routing distribution into a rigid step function. This constraint introduces a brittle boundary where highly competitive experts are arbitrarily separated into full-supervision and zero-feedback zones based on minor score fluctuations. To address this issue, we propose Elastic Expert Routing, which stochastically samples the active expert budget from a localized discrete distribution centered at $k$. Over multiple training iterations, this mechanism softens the sharp threshold into a gradual probability distribution. Because the sampling neighborhood remains symmetric, this approach matches the expected computational cost of deterministic training, while preserving the inference budget. Extensive experiments demonstrate the efficacy of our method on both supervised fine-tuning and from-scratch pretraining settings. During supervised fine-tuning, elastic routing improves downstream macro-averages on OLMoE-1B-7B and Qwen3-30B-A3B by $+0.84$ and $+2.02$ points, respectively. In addition, in from-scratch pretraining, it outperforms the static top-$k$ baseline by $1.6$ points on average across downstream tasks.