论文

线性 ViT 初始化新方法:复制 MLP 权重、蒸馏注意力

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

精选理由

想让线性 ViT 免去从头预训练?这篇告诉你直接复制 MLP 权重再蒸馏注意力就能追平 Softmax 版本,方法很实用。

研究针对线性 Vision Transformers 替换 Softmax ViT 注意力后需从头预训练的问题,提出 Softmax-to-linear 权重迁移方案。实验发现注意力权重是算子特定的,直接复制甚至可能不如随机初始化,但通过蒸馏可恢复其 token 路由行为。MLP 权重与算子无关,直接复制即可获得预训练权重的绝大部分收益。两者结合后,线性 ViT 能追平甚至超越 Softmax 版本,且结论在多种线性 ViT 变体、不同模型规模和数据集上一致成立。

原文 · arXiv cs.LG

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and effectively still remains unclear. In this work, we explicitly ask: given that most foundation ViTs are built on the mainstream Softmax attention, can linear ViTs benefit from their pre-trained weights? Recent works on Attention Transfer show that attention is the effective transferable component between Softmax ViTs, suggesting attention alone suffices for such reuse. However, we find the opposite for Softmax-to-linear transfer. The attention weights are operator-specific: copying them barely helps, and is sometimes even worse than random initialization. Instead, the attention's token routing behavior can be recovered through distillation with a proper loss design, letting linear ViTs reduce the gap and even match Softmax ones. In contrast, the MLP weights, which carry the learned representation, are operator-agnostic: they can be transferred by simple direct copying, which already carries most of the benefit of the pre-trained weights. Thus, copying MLPs can serve as an effective foundation for Softmax-to-linear transfer: paired with the distilled attention, linear ViTs eventually close the remaining gap and even surpass Softmax ones. These findings hold consistently across various linear ViT variants, different model sizes, and diverse datasets. We hope this study deepens the understanding of reusing pre-trained weights across attention operators: copy what stays the same and distill what differs, to recover the benefit across the Softmax-to-linear boundary.