强化学习扩展无需更多训练
Does Scaling Reinforcement Learning Really Require More Training?
SURGE方法让强化学习训练历史发挥更大价值,数学和编程任务准确率超越训练曲线最高点。
研究人员提出策略空间扩展方法SURGE,通过融合同一强化学习运行中的两个检查点来提升模型性能。该方法在DeepSeek AIME24基准上达到54.17%准确率,超过原生最高50.83%;在OLMo HumanEval+上达到83.7%,超过82.8%。SURGE在保留锚点更新主要成分的同时融入捐赠者的互补成分,无需延长训练或增加推理计算量。
Does Scaling Reinforcement Learning Really Require More Training?
Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL history, without extending training or increasing per-query inference computation. We instantiate it with SURGE (Scaling Up RL Gradient-free via Eigenspace fusion). SURGE combines two checkpoints from the same RL run: a high-accuracy anchor and a competitive donor that generates shorter responses. It expresses both checkpoints as changes from their shared initialization, then spectrally decomposes the anchor's update to retain its dominant component and incorporate the donor's complementary component. With a fixed target for how much of the anchor update to retain, SURGE determines the block size from the weights without testing candidate policies. We evaluate two 1.5B mathematical-reasoning histories, DeepSeek and Nemotron, and one 7B coding history, OLMo. SURGE improves benchmark-average accuracy over both input checkpoints while using fewer reasoning tokens than the anchor. It reaches 54.17% on DeepSeek AIME24 against a measured native maximum of 50.83%, and 83.7% on OLMo HumanEval+ against 82.8%. These gains exceed the observed training curves. Geometric controls support the importance of RL-update structure beyond weight displacement or token reduction alone. Each constructed model runs as a single policy. Our findings identify stored RL history as a reusable scaling resource: the capability available from a training run need not end at its best checkpoint.