论文

MoRA:通过路由器偏置学习和专家近似进行MoE剪枝

MoRA: MoE Pruning via Router Bias Learning and Expert Approximation

精选理由

MoRA解决了MoE模型部署内存问题,通过路由偏置学习和专家近似机制,在减少专家数量的同时保持模型性能。

MoRA是一种新的混合专家模型剪枝框架,在Qwen3-30B-A3B、DeepSeek-V2-Lite和Moonlight-16B-A3B模型上测试。该框架通过可学习的路由器偏置和专家近似机制,每层移除25%和50%的路由专家。在九个零样本基准测试中,MoRA超越了最先进的剪枝算法。

原文 · arXiv: DeepSeek

MoRA: MoE Pruning via Router Bias Learning and Expert Approximation

Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts. However, existing pruning methods either use expert ranking criteria that are not well aligned with model performance or rely on effective expert subset searching that is computationally expensive. Moreover, these methods typically overlook the routing-behavior redundancy among the retained experts. In this paper, we propose MoE Pruning via Router Bias Learning and Expert Approximation (MoRA), a framework for structured MoE expert pruning. We introduce a learnable router bias for each expert and optimize these biases by minimizing the language-modeling loss and a routing-diversity regularizer. The learned router biases sharpen the routing probability distributions to identify experts critical to model performance while encouraging the selection of experts with diverse routing preferences. In addition, we introduce an expert approximation mechanism as a post-pruning enhancement. It leverages the remaining experts to approximate the outputs of pruned experts by affine transformation, further improving the performance of the pruned model. We evaluate MoRA on Qwen3-30B-A3B, DeepSeek-V2-Lite, and Moonlight-16B-A3B, removing 25\% and 50\% of the routed experts in each MoE layer. Extensive experiments on nine zero-shot benchmarks show that MoRA outperforms state-of-the-art pruning algorithms. Our code will be released.