MeqMuon:矩阵均衡化 Muon 优化器改进 LLM 预训练
MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining
训练大模型的朋友可以看看这个 Muon 改进版,行列一起均衡还省优化器显存,收敛比 AdamW 更好。
这篇论文提出矩阵均衡化 Muon(MeqMuon)优化器,用于 LLM 预训练。相比只做行归一化的 Muon 变体,MeqMuon 同时均衡更新矩阵的行和列幅值,无需手动调节即可适配不同失衡模式。它还省去 AdamW 二阶矩估计的存储,降低优化器状态显存占用。实验显示其收敛表现优于 AdamW、Muon 等基线。
MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining
The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose an improved Muon optimizer, called \underline{m}atrix-\underline{eq}uilibrating Muon~(MeqMuon), for LLM pretraining. MeqMuon balances both row and column magnitudes through normalization that can be automatically tailored to different imbalance patterns without manual intervention. Moreover, MeqMuon eliminates the need to store AdamW's second-moment estimates, reducing optimizer-state memory usage. Empirical results demonstrate that MeqMuon achieves better convergence performance than AdamW, Muon, and other baselines in LLM pretraining.