论文

ITC-MoE:MoE扩散语言模型的压缩框架

ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models

精选理由

清华团队提出ITC-MoE框架,能压缩MoE扩散模型30%参数同时保持96%准确率,速度提升7倍。

ITC-MoE是一种针对MoE扩散语言模型的压缩框架,包含IATC和TCR两个组件。该框架在SDAR-30B-A3B-Chat-b32模型上测试,30%压缩预算下MultiArith准确率达96.33%,端到端速度提升最高7.22倍。ITC-MoE解决了跨模式非均匀冗余和标记级利用变化的问题。

原文 · arXiv cs.AI

ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models

Mixture-of-Experts (MoE) Diffusion Language Models (DLMs) offer flexible parallel decoding and increased model capacity, but their large number of expert parameters incurs substantial computation and storage costs. Existing low-rank MoE compression methods largely rely on static factorization and fixed rank allocation, which overlook the distinctive properties of MoE DLMs. Specifically, we identify two properties: cross-mode non-uniform redundancy, where parameter redundancy and sensitivity to rank truncation vary across the input, output, and expert modes, and token-wise utilization variation, where hot and cold tokens exhibit distinct spectral characteristics and expert activation patterns. To address these challenges, we propose ITC-MoE, an Importance-guided Token-aware Compression framework for MoE DLMs. ITC-MoE consists of two complementary components. First, Importance-guided Adaptive Tucker Compression (IATC) incorporates activation and gradient importance into expert weight transformation, jointly factorizes expert weights across multiple modes, and adaptively allocates ranks under a fixed parameter budget. Second, Token-aware Compensation and Routing (TCR) applies lightweight low-rank compensation to compression-sensitive hot tokens and restricts the candidate expert set for cold tokens with concentrated routing patterns. By jointly adapting compression capacity and inference execution to both parameter redundancy and token-wise variation, ITC-MoE substantially reduces the computation and storage costs of MoE DLMs while preserving their generation quality. For example, on SDAR-30B-A3B-Chat-b32, ITC-MoE maintains an accuracy of 96.33% on MultiArith under a 30% compression budget, while achieving up to a 7.22x end-to-end speedup. The code is publicly available at https://github.com/lianjunl13-sudo/ITC-MoE.