Off-Policy Merging 持续学习方法超越 On-Policy 自蒸馏
Off-Policy Merging Beats On-Policy Self-Distillation for Continual Learning
做模型持续训练的必看:不用贵的 on-policy 采样,靠 grafting 三步就能加新能力还不忘旧本事,比 SFT 和 OPSD 都强。
一篇 arXiv 论文提出名为 grafting 的持续学习配方,解决 SFT 在后训练模型上造成灾难性遗忘的问题。该方法分三步:在更早的 donor checkpoint(甚至预训练结束前)上学习权重更新,再将更新以模型合并方式缩放后应用,并在数据分布差异大时屏蔽最敏感的更新方向。实验涵盖专家轨迹蒸馏、STaR 与 Pedagogical RL 自我改进、预训练截止后知识注入三类场景,grafting 在新任务与旧任务表现上均 Pareto 优于 SFT 和 OPSD。论文结论挑战了 on-policy 训练是 RL 训练模型持续学习必要条件这一传统观点。
Off-Policy Merging Beats On-Policy Self-Distillation for Continual Learning
A long-standing goal of AI is a model that can continually learn and improve itself. On post-trained models, supervised finetuning (SFT) on new data often causes poor generalization and catastrophic forgetting. As such, the conventional wisdom is that on-policy training is a prerequisite for continual learning. In practice, however, data containing new knowledge or capabilities are often off-policy. While methods such as on-policy self-distillation (OPSD) try to bridge this gap by converting off-policy data into on-policy signal, they have been shown to cause reasoning collapse. In this paper, we show that off-policy merging beats OPSD for continual learning. We first show that SFT learns a useful signal from new data, but naively applying its update interferes with existing capabilities. We reduce this interference with a simple recipe we term grafting, which changes where the update is learned and how it is applied: (1) learning the update on an earlier donor checkpoint, ideally even before the end of pretraining, and applying the weight update to the post-trained model; (2) scaling the weight update, equivalent to a form of model merging; and (3) optionally, masking the most sensitive update directions when the new data distribution is far from the post-trained model. Across continual learning settings including (1) distilling from expert traces, (2) self-improvement with STaR and Pedagogical RL, and (3) injecting knowledge after pretraining cutoff, grafting Pareto-dominates both SFT and OPSD in new-task and old-task performance, while avoiding expensive on-policy sampling. Therefore, our work challenges on-policy training as a necessity for continual learning on RL-trained models.