论文

多教师在线蒸馏研究:从梯度到能力

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

精选理由

研究Qwen3如何整合多个RL训练教师,发现梯度传递和参数更新的关键机制。

该研究探索了Qwen3-1.7B模型如何整合四个领域教师模型的优势。研究发现损失平均隐式加权响应,Adam的一阶矩减少参数更新差异,BF16舍入隐藏微小变化。在数学任务上,响应平均比采样令牌策略梯度准确度高2.6点,但全局令牌平均则低2.1点。

原文 · arXiv cs.LG

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97\% of FP32 master weights differ from initialization, but only 7--11\% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.