论文

NorMuon行归一化研究

The Row Normalization Puzzle in Muon

精选理由

NorMuon论文解析了行归一化在理论保证与实践表现间的矛盾,对理解LLM优化很有价值。

该论文研究了行归一化对Muon模型的影响,揭示了NorMuon最坏情况保证与实际表现之间的差距。研究显示行归一化在算子范数几何下引入了维度相关因子,即使使用精确极计算和固定动量参数仍然存在。实验表明NorMuon在合成问题上比Muon慢,但在LLM预训练中表现更优。

原文 · arXiv cs.LG

The Row Normalization Puzzle in Muon

This paper examines how row-wise renormalization affects Muon, focusing on the gap between NorMuon's worst-case guarantees and its practical performance (Li et al.). Despite its growing adoption and promising performance in large language model (LLM) pretraining, NorMuon's worst-case guarantees remain poorly understood. One fundamental question is: Does row normalization yield provable convergence gains, potentially through its interaction with approximate polar computation and exponential moving-average momentum? Our results show that row normalization introduces a dimension-dependent factor in the worst-case iteration complexity under the operator-norm geometry, which persists even with exact polar computation and any fixed momentum parameters. Indeed, we establish an algorithm-dependent lower bound and a matching upper bound in deterministic settings, and extend our upper bound analysis to stochastic settings. Both upper-bound analyses allow approximate polar computation. Experiments show that NorMuon is slower than Muon on synthetic problems inspired by our worst-case construction, yet outperforms Muon in LLM pretraining. These findings sharpen the puzzle of why row normalization helps in practice and complement the recent findings of Dewulf et al.