CIPHER-MoE:万亿参数 MoE 训练负载均衡新方法
CIPHER-MoE: Balancing Efficiency and Routing Fidelity in Trillion-Scale MoE Training
训练万亿参数 MoE 模型的负载不均问题有了新解法,DeepSeek-V4-Pro 上实测训练加速最高 1.94 倍,还不用加硬件,做训练的可以看看。
arXiv 论文提出 CIPHER-MoE,通过亲和度感知的 Expert-to-Token 过滤和显式容量控制来缓解 MoE 训练中专家负载不均的问题,且不改动路由器 token 侧的 Top-K 选择。该方法在 DeepSeek-V4-Pro 等大规模 MoE 模型上完成评估,Top-1 专家负载最高减少 64.9 个百分点,训练加速 1.10 倍到 1.94 倍。与依赖复杂并行策略或资源再分配的系统级方案不同,它不需要额外硬件资源或复杂的运行时设计,同时保持训练质量。源代码即将开源。
CIPHER-MoE: Balancing Efficiency and Routing Fidelity in Trillion-Scale MoE Training
Mixture-of-Experts (MoE) has been widely adopted in recent large language model (LLM) architectures. However, scaling up MoE in LLM training introduces system-level challenges on training, where non-uniform token routing can lead to highly imbalanced workloads across experts and devices, further destabilizing the training process. With trillion-scale LLMs, imbalanced expert workloads further amplify the resource cost of MoE training, resulting in degraded training efficiency and hardware utilization for underloaded experts, while hot experts require additional resources to accommodate excessive workloads. Recent studies address imbalanced MoE training through intricate parallelism strategies or resource reallocation. However, these system-level approaches often introduce additional resource requirements and considerable orchestration complexity, which become increasingly difficult to afford when training trillion-parameter LLMs under constrained computational resources. This work introduces CIPHER-MoE, which mitigates MoE workload imbalance while keeping the router's token-side Top-K selection unchanged. CIPHER-MoE applies affinity-aware Expert-to-Token filtering with explicit capacity control to reduce hotspot expert workloads without additional hardware resources or complex runtime design. The proposed method has been evaluated on large-scale MoE models, including DeepSeek-V4-Pro, showing up to 64.9 percentage points Top-1 expert workload reduction and 1.10$\times$-1.94$\times$ training acceleration, while preserving the training quality. The source code will be released soon.