论文

Pivot-SD:用自蒸馏提升掩码扩散语言模型推理能力

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

精选理由

只拿 200 道题微调,就让 LLaDA-8B-Instruct 的数学和代码成绩跑赢 SFT 和扩散 RL,方法还挺省算力,做扩散模型训练的可以看看。

arXiv 论文提出 Pivot-SD,一个针对掩码扩散语言模型(dLM)的离线自蒸馏框架。该方法用信息增益指标识别去噪过程中的关键 token 承诺(pivots),只对这些高影响位置做监督:成功轨迹用交叉熵训练,失败轨迹的关键位置用 targeted unlikelihood。仅用 200 道题、每题 4 次 rollout,Pivot-SD 让 LLaDA-8B-Instruct 在数学和代码基准上超过全序列 SFT 和预算对齐的扩散 RL 基线。

原文 · arXiv cs.LG

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an information-gain metric measuring uncertainty reduction over the remaining masked positions. Pivots from successful trajectories are trained with cross-entropy, and pivots from failed trajectories with targeted unlikelihood, leaving the rest of the failed trajectory untouched. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.