SplitMoE突破视频扩散模型均匀性陷阱
Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
清华团队提出SplitMoE,解决视频扩散模型均匀性陷阱,生成质量更优。
SplitMoE是一种新型稀疏架构,专为视频数据设计。该模型将专家池分为语义专家和通用专家两类,在相同激活参数预算下,SplitMoE在收敛速度、路由连贯性和视频生成质量上均优于传统负载均衡MoE。研究团队在标准基准测试中验证了SplitMoE的有效性。
Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.