研究用强化学习视角解释 On-policy Distillation 崩溃成因
Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective
把 OPD 训练崩溃讲透了,还给了屏蔽坏响应和 SFT 初始化两个可上手的缓解办法,做蒸馏的人快看。
arXiv 论文分析了 on-policy distillation(OPD)中性能提升与生成长文本、重复内容崩溃这两种相反结果的形成机制。实验显示 OPD 在不扩展学生模型能力的前提下提升表现,当隐式奖励与质量错位时会触发 reward hacking,放大冗长重复的输出。作者提出训练时屏蔽不健康响应和采用 SFT 初始化两种方法缓解崩溃,代码已在 GitHub 开源。
Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective
On-policy distillation (OPD) has become an important approach to language model post-training. However, despite its performance gains, OPD can also collapse into excessively long and repetitive generation, and the mechanism underlying these divergent outcomes remains poorly understood. We explain these outcomes through a reinforcement learning perspective: the teacher implicitly rewards student behaviors, even those it rarely exhibits itself. From this perspective, our experiments show that OPD improves performance without expanding the student's capabilities. When the implicit reward model is reliable, OPD makes correct responses easier to sample. In contrast, when the preference misaligns with quality, reward hacking happens: the implicit reward model amplifies overlong, repetitive student rollouts, even though it rarely generates such text itself. Guided by this diagnosis, we find that masking unhealthy responses during training and using SFT initialization can each effectively mitigate the collapse. Together, these findings show that OPD amplifies student behaviors favored by the teacher's implicit feedback, shifting the focus from how well the teacher generates to how reliably it evaluates student rollouts. Our code is available at https://github.com/HancCui/opd_hacking.