论文精选

OPPD 蒸馏法让模型单次生成即达 power sampling 64 候选的效果

Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution

精选理由

一篇有意思的论文:不靠每次采样几十个候选,而是把 power sampling 的锐化直接蒸进模型,单次生成就超过 64 候选采样,数学和代码基准都有实测数字。

OPPD(on-policy power distillation)用顺序蒙特卡洛采样器,让被训练模型生成候选、由冻结教师的 power 分布加权,再做最大似然更新,从而把 power sampling 的锐化效果直接训进模型。在 MATH500 上单次生成准确率较未训练模型提升最多 23.0 点,GSM8K 提升 27.3 点;单次生成比已发表的 64 候选 power sampling 还高 2.4 和 3.5 点,并恢复了 16 候选带来增益的 94%。与同预算的 GRPO 相比,OPPD 在 MATH500、GSM8K、AIME 上分别高 3.8、4.0、5.4 点,且无需参考答案;在 GRPO 之后叠加 OPPD 可再加最多 9.3 点。仅在数学上训练的 OPPD 还能把 HumanEval 准确率提高最多 5.3 点,且在已用可验证奖励训练过的模型上仍能再加 4.4 点(MATH500)。

原文 · arXiv cs.AI

Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution

A language model can give a correct answer more probability than any single incorrect answer and still usually sample an incorrect one, because the incorrect answers together hold more probability. The power distribution raises each complete answer's probability to a power above one and renormalizes, shifting probability toward answers the model finds most likely (sharpening). Sampling from it improves reasoning without changing parameters, but needs many scored candidates per query. We show that a model can instead be trained to produce such answers in one generation. On-policy power distillation (OPPD) runs a sequential Monte Carlo sampler in which the model being trained generates candidates and a frozen teacher's power distribution weights them; the same probabilities weight each answer in a maximum-likelihood update. Training raises single-generation accuracy by up to 23.0 points on MATH500 and 27.3 on GSM8K over the untrained model at the same temperature, and one generation scores 2.4 and 3.5 points above published power sampling with 64 candidates, recovering 94 percent of the gain that 16 candidates give the untrained model. For context, against GRPO trained with verified rewards from the same checkpoint and budget, OPPD scores 3.8, 4.0 and 5.4 points higher on MATH500, GSM8K and AIME using no reference answers; the two are complementary, and OPPD applied after GRPO adds up to 9.3 points. Trained only on mathematics, OPPD raises HumanEval accuracy by up to 5.3 points. One loss coefficient moves the sharpening exponent the model absorbs between 1.19 and 2.02, against 1.14 for ordinary on-policy distillation, and it rises mostly on the model's own answers. Gains hold across model families and sizes, including a model already trained with verified rewards, where lowering the temperature gives nothing and OPPD adds 4.4 points on MATH500. Code: https://github.com/ArminAzizi98/OPPD.