论文精选

多令牌预测头后训练方法研究

Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads

精选理由

DeepSeek团队发现,只需少量后训练就能让Qwen3-8B达到与MiMo-7B相当的推理速度,比原方法节省千倍训练资源。

研究人员提出了一种针对冻结推理模型的后训练方法,在Qwen3-8B模型上使用约25亿个token进行训练,达到了与MiMo-7B联合预训练相当的数学、编程和知识基准性能。该方法比MiMo-7B的联合预训练减少了10^3-10^4倍的MTP训练token。研究还提出了链式感知的验证规则放松,在保持任务准确度的同时提升了12-16%的加速效果。此外,自适应控制器能动态选择MTP头数量,在固定最大MTP草稿长度下恢复了11-14%的加速损失。

原文 · arXiv: DeepSeek

Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads

Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (MiMo-7B, DeepSeek-V3, Qwen3) trains its heads jointly with the backbone over the full pretraining run of tens of trillions of tokens, thus setting the drafter quality at pretraining time. We ask whether a lightweight post-training pass on target-generated chain-of-thought is enough to reach the same expected throughput speedup on a frozen reasoning model, and study how a serving-time system built on such a checkpoint can be optimized. We present three findings. 1) On a frozen Qwen3-8B with $K{=}3$ chained MTP heads, we show that a post-training recipe with plain cross-entropy on $\approx\!2.5$B tokens reaches or exceeds the expected speedup of jointly trained MiMo-7B on math, coding and knowledge benchmarks. Our post-training recipe utilizes $10^3$-$10^4\times$ less MTP-training tokens as compared with joint pre-training of MiMO-7B MTP baseline. 2) We propose a chain-aware relaxation of draft token verification rule that allows a bounded drift from backbone language model token distribution. We show that this relaxation lifts expected speedups by $+12$ to $+16\%$ per benchmark while preserving task accuracy. 3) We propose an adaptive controller that dynamically chooses the number of MTP heads to be engaged at inference time and demonstrate recovery of upto $11$--$14\%$ loss in speedup using fixed maximum MTP draft length.