BRANCH-MoE 提出树状路由结构,改进 MoE 专家负载均衡
BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models
这篇论文把 MoE 的专家排成二叉树来做路由,不用额外的负载均衡损失,还能按树前缀分设备减少通信,做 MoE 训练的可以看看。
arXiv 论文提出 BRANCH-MoE,将 E 个专家放在深度为 log2 E 的二叉决策树叶节点上,分支概率以到达流量的加权均分为中心,用指数移动平均估计。该机制无需辅助负载均衡损失即可避免路由质量坍缩,论文在线性节点映射和对数凹到达分布下给出了证明。在 Criteo、Forest Covertype、HIGGS、YearPredictionMSD 四个数据集上,使用 E=16、top-4 路由、五个随机种子,与 Switch softmax、DeepSeek-V3 dynamic-bias、Skywork logit-normalized 和 hash 路由对比。结果显示分层路由在保持任务质量和均衡利用率的同时,可按树前缀分配专家到设备,降低跨设备通信量。
BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models
Mixture-of-experts (MoE) layers increase model capacity without a proportional increase in per-example computation. However, conventional flat routers can yield imbalanced expert utilization and treat experts as an unstructured collection, whose indices carry no topological meaning. We introduce {\bf BRANCH-MoE}, a routing architecture that places \(E\) experts at the leaves of a binary decision tree of depth \(\log_2 E\). At each internal node the branching probability is centered on the arrival-weighted mean score of the traffic reaching that node. This mean is estimated using an exponential moving average, which promotes utilization of both child subtrees without an auxiliary load-balancing loss. We show that this moving-average estimate admits an explicit noise-lag trade-off. We prove that for linear node maps and log-concave arrival distributions, this mechanism prevents routing-mass collapse. We further establish that, under a frozen router, an expert's execution frequency controls its stochastic-gradient convergence rate, and that confident decisions near the root bound cross-device communication when experts are assigned to devices by tree prefix. We evaluate BRANCH-MoE against Switch softmax, DeepSeek-V3 dynamic-bias, Skywork logit-normalized, and deterministic hash routing on Criteo click-through-rate prediction, Forest Covertype, HIGGS, and YearPredictionMSD, using \(E=16\), top-\(4\) routing, and five random seeds. Our results show that hierarchical routing can preserve task quality and balanced utilization while inducing a topology that supports localized expert co-activation and reduced communication.