论文精选

论文提出 EDR 目标函数优化并行投机解码草稿模型训练

Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds

精选理由

一篇加速 LLM 推理的论文,把投机解码建模成马尔可夫过程,微调 DSpark 和 DFly 后在九个基准上超过原有训练方法。

arXiv 论文将投机解码表示为马尔可夫奖励过程,提出 Expected Decoding Rounds(EDR)目标,直接等于期望解码轮数。EDR 无需辅助超参数,并推导出基于 target 模型 rollout 的精确时序差分梯度。该框架还提供离线评估器,可在共享 rollout 上成对比较草稿模型而无需实际运行投机解码。在 DSpark 和 DFly 两个草稿模型上微调后,EDR 在数学推理、代码生成和聊天共九个基准上均优于现有训练目标。

原文 · arXiv cs.AI

Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds

Speculative decoding accelerates large language model inference by using a low-cost draft model to propose tokens that the full-size target model verifies in parallel. Parallel and semi-autoregressive (semi- AR) drafters improve drafting efficiency by proposing an entire block in a single forward pass, but training them raises a new difficulty: the draft distribution for a given position depends on where the decoding round starts, and where rounds start depends on how many tokens earlier rounds accepted. Existing training objectives typically rely on block-local surrogates that ignore this cross-round coupling, and therefore do not directly optimize the global decoding efficiency. In this work, we develop a theoretical framework for training and evaluating these drafters by representing speculative decoding as a Markov reward process. This formulation yields the Expected Decoding Rounds (EDR) objective, which weights local rejection costs by state occupancies and exactly equals the expected number of decoding rounds. Unlike prior surrogate objectives, EDR introduces no auxiliary hyperparameters. We then derive an exact temporal-difference gradient that supports unbiased stochastic optimization from target-model rollouts. The same framework also yields an exact offline evaluator for round counts, enabling paired drafter comparisons on shared target rollouts without running speculative decoding. Finetuning two state-of-the- art drafters, DSpark and DFly, with EDR consistently improves mean accepted length and outperforms existing training objectives across nine benchmarks spanning math reasoning, code generation, and chat.