论文

Hessian 特征值聚类之谜:对称性破缺理论给出统一解释

Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking

精选理由

训练好的模型 Hessian 谱总长那样?这篇论文用对称性破缺把聚簇和离群值一次讲清楚,还覆盖 transformer 和 NTK。

这篇论文解释了深度学习中 Hessian 谱的常见模式:特征值分为靠近零的主体簇和少数离群值。作者将其归因于对邻近高对称参考构型的偏离,该参考构型的 Hessian 具有超出权重对称性的丰富不变性。论文对三层 ReLU 网络做了详细分析,并应用到卷积、图、transformer 模型和 NTK,同样的机制也出现在逐层 Hessian 和 Gauss-Newton 矩阵中。

原文 · arXiv cs.LG

Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking

Hessian spectra at trained models in deep learning exhibit a persistent pattern: eigenvalues organize into distinct clusters, including a large bulk near zero and a few isolated outliers. This paper shows that a natural account of these spectral phenomena emerges when the original setting is understood as a departure from a nearby, otherwise hidden, highly symmetric reference. Modifications, including changes to the architecture, data distribution, or parameter metric, expose a nearby reference configuration whose Hessian exhibits rich invariances-ones not accounted for by weight symmetries. There, symmetry enables a precise description of the spectra, forcing high-dimensional kernels and eigenvalues of large multiplicity. Returning to the original configuration breaks the Hessian symmetry and thereby produces the observed hierarchy of clusters and outliers. The framework is developed in some generality, with a detailed analysis of three-layer ReLU networks and applications to convolutional, graph, and transformer models, as well as to the NTK. The same mechanism is further shown to yield analogous spectral structures in layerwise Hessians and the Gauss-Newton matrix.