论文

SoftServe:深度学习可扩展拟牛顿方法

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

精选理由

研究人员推出SoftServe拟牛顿方法,解决深度学习中的非凸性和大规模参数问题,在病态问题上表现优于现有基线。

SoftServe是一种新型拟牛顿方法家族,专为克服深度学习中的非凸性和大规模参数问题设计。该方法从Berglund等人(2025)的变分目标中推导出正定曲率估计,即使在存在负曲率的情况下也能工作。SoftServe包含对角和Kronecker分解变体,通过构造保持正定性,可扩展至大规模神经网络。在136M参数物理信息扩散模型等严重病态问题上,SoftServe实现了比Adam、Muon和SOAP等基线更低的损失。

原文 · arXiv cs.LG

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

Quasi-Newton (QN) methods have long been among the most effective methods for large-scale unconstrained convex optimization. Two obstacles have limited their use in deep learning: non-convexity and enormous parameter sizes. We introduce SoftServe, a family of QN methods designed to overcome these obstacles without line searches or ad hoc curvature corrections. SoftServe derives positivedefinite curvature estimates from the variational objective of Berglund et al. (2025), even in the presence of negative curvature. We develop diagonal and Kroneckerfactored variants that preserve positive definiteness by construction and scale to massive neural networks. Finally, SoftServe relies on the stable coupled Newton-Schulz iteration for the required matrix operations, replacing costly matrix decompositions with GPU-friendly matrix multiplications. SoftServe excels on problems that are severely ill-conditioned, including tasks such as recurrent networks, deep autoencoders, physics-informed neural networks, and a 136M-parameter physics-informed diffusion model, often achieving lower losses than established baselines including Adam, Muon, and SOAP.