论文

论文提出 horizon loss:一行改动让分类训练超过交叉熵

Planning to Learn

精选理由

有人把交叉熵重新解读成策略梯度,然后推出一行代码就能替换的 horizon loss,在 ResNet-50 和 ViT-S/16 上实测涨点,做分类训练的可以试试。

论文从策略梯度视角重新审视分类训练,指出交叉熵优于精确策略梯度的原因是后者的梯度过于短视,只看当前收益而忽略对后续学习的影响。作者提出 horizon loss,按剩余学习量截断总误差,只需一行代码改动。在 MNIST 以及 ImageNet 上的 ResNet-50、ResNet-101 和 ViT-S/16 实验,horizon loss 在平坦学习率下的 top-1 准确率超过交叉熵,且标签噪声越大提升越明显。

原文 · arXiv cs.LG

Planning to Learn

Policy-gradient methods are central to modern reinforcement learning, including LLM post-training. When they struggle, the usual suspects are exploration, credit assignment and action-sampling noise. Classification has none of them. A classifier is a policy whose expected reward, its \emph{expected accuracy}, is the probability it assigns to the correct label, and because that label is known, the policy gradient is exact and smooth. Yet exact policy gradient loses to cross-entropy, even on expected accuracy. The exact gradient is myopic: it values an update only by what it buys now, but each update also sets where the next one starts, so an update's value depends on how much learning remains. Viewed this way, cross-entropy is patient accuracy, the total error an example would pay if its log-odds rose at unit speed forever, while exact policy gradient is the zero-horizon limit. Truncating this total at the learning that remains yields the horizon loss, a one-line change that moves from cross-entropy toward exact policy gradient as training runs out. In a simple allocation model, it provably escapes the trap that catches each endpoint. On MNIST and on ImageNet with ResNet-50, ResNet-101 and ViT-S/16, the horizon loss improves top-1 accuracy over cross-entropy at a flat learning rate, and the gain grows with label noise.