MC-Sparse:训练-free 稀疏注意力框架缩小 DiT 稠密稀疏差距
MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
一篇实用论文:MC-Sparse 不用训练就能把视频和 3D 生成的去噪提速 1.8 到 2.3 倍,还指出了以往稀疏注意力掉质量的三个原因。
arXiv 论文提出 MC-Sparse(Meta-Cached Sparse Attention),一种无需训练的稀疏注意力框架,用于加速扩散 Transformer 的视频与 3D 资产生成。论文先用 oracle 对比把质量退化归因到 token 分组约束、交互选择不准和丢弃 token 造成的注意力损失三个来源。MC-Sparse 通过缓存查询分组、按精确注意力概率选出的 KV 索引以及稠密-稀疏输出残差,在后续去噪步骤复用。在 Minimax-H3-Base 上相对稠密注意力获得 1.80 倍去噪加速,3D 资产生成获得 2.32 倍加速,质量损失可忽略。
MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing methods can degrade generation quality and fidelity at high sparsity levels. Through controlled oracle comparisons, we trace this degradation to three sources: constraints imposed by token grouping, inaccurate interaction selection, and the attention contributions lost when tokens are discarded. Guided by this analysis, we propose Meta-Cached Sparse Attention (MC-Sparse), a training-free framework that selects individual key-value (KV) tokens while organizing similar queries into tile-aligned groups for efficient GPU execution. MC-Sparse caches metadata comprising query groups, KV indices selected using exact attention probabilities, and residuals between dense and sparse attention outputs, and reuses them across subsequent denoising steps. Across video and 3D generation models, MC-Sparse achieves higher fidelity to dense-attention outputs and larger denoising speedups than existing sparse-attention baselines, without visible quality degradation. Relative to dense attention, it delivers a $1.80\times$ denoising speedup on Minimax-H3-Base and a $2.32\times$ speedup on 3D asset generation, both with negligible quality loss.