论文精选

Activation-Conditioned Self-Distillation研究发布

Activation-Conditioned Self-Distillation

精选理由

ACSD让模型自己当老师,数学准确率提升近3%,还能跨数据集复用方向向量。

ACSD方法通过对比正确轨迹与所有轨迹的激活提取转向向量,实现无需参考文本的教师参数更新的自蒸馏。在五个模型上,ACSD在四个数学基准测试中取得最高平均准确率。DeepSeek-R1-0528-Qwen3-8B模型在ACSD下数学准确率达71.9%,LiveCodeBench v6 pass@12达70.9%,优于基准线69.0%和66.3%。

原文 · arXiv: DeepSeek

Activation-Conditioned Self-Distillation

On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning. Providing privileged information does not by itself ensure effective token-level supervision throughout long responses. We introduce Activation-Conditioned Self-Distillation (ACSD), which extracts a steering vector by contrasting activations of self-generated trajectories that reach verified correct answers within a generation budget with those of all remaining trajectories. A frozen copy of the base model applies this vector at each prediction position, and the student learns from its next-token distributions on student-generated prefixes. Outcome verification is used for direction construction and calibration; distillation requires neither problem-specific reference text nor teacher parameter updates. The distilled student is used alone at inference. On each of five models, ACSD achieves the highest mean accuracy over four mathematical benchmarks among the evaluated methods. On DeepSeek-R1-0528-Qwen3-8B, mean mathematical accuracy reaches 71.9\% and LiveCodeBench v6 pass@12 reaches 70.9\%, compared with 69.0\% and 66.3\% for the reference-conditioned OPSD baseline. Contrasts among correct trajectories also support distillation, and extracted directions can be reused across mathematical training datasets. On fixed student trajectories, ACSD maintains more stable late-position logit-update magnitudes than OPSD.