论文

NS-Attn:用 Newton-Schulz 变换注意力输出提升 ViT 准确率

NS-ATTENTION: Newton-Schulz Transformations of Attention Outputs in Vision Transformers

精选理由

把 Muon 里那套 Newton-Schulz 迭代搬到注意力输出上,不加参数,ViT 和 Swin 在 CIFAR 上 12 组对比全涨,代价是推理变慢一点。

一篇 arXiv 论文提出 NS-Attn.,把 Muon 优化器中使用的 Newton-Schulz 迭代直接应用到 Transformer 注意力头的输出上,这是一种无需参数的变换。做法是将每个头的输出排成特征-令牌矩阵,按 Frobenius 范数归一化后执行一次有限 NS 多项式步骤,以降低谱集中度、提高有效秩。在 CIFAR-10 和 CIFAR-100 上对 ViT 和 Swin 的 12 组匹配种子对比中,最终轮准确率全部提升,平均增益 0.25–0.83 个百分点。ViT 消融显示一次迭代优于两次,代价是额外的推理延迟。

原文 · arXiv cs.LG

NS-ATTENTION: Newton-Schulz Transformations of Attention Outputs in Vision Transformers

Newton-Schulz (NS) iteration has recently been used in the Muon optimizer to transform update matrices during the training of large language models. Motivated by its spectral effect, we investigate applying NS directly to Transformer attention representations. We introduce Newton-Schulz Attention (NS-Attn.), a parameter-free transformation applied to the output of each attention head. Each head output is arranged as a feature-by-token matrix and normalized by its Frobenius norm. We then apply a finite NS polynomial step and restore the original norm. The objective is to reduce spectral concentration and increase effective rank before standard head merging and output projection. Across ViT and Swin on CIFAR-10 and CIFAR-100, NS-Attn. improves final-epoch accuracy in all 12 matched-seed comparisons, with mean gains of 0.25--0.83 percentage points. ViT ablations show higher mean accuracy with one iteration than with two. Spectral analysis further shows reduced leading-eigenvalue concentration and increased effective rank. These gains incur additional inference latency.