从推理链哪一步重启更划算?arXiv 论文量化扩展位置的选择收益
Expanding LLM Reasoning
论文发现重启推理链的位置本身就有讲究,选对位置在 MATH-500 上能用更少算力超过 self-consistency,做推理成本优化的人可以看。
arXiv 论文 2610.05584 研究:额外推理算力不该只用来采样更多推理链,还应选好从已存步骤的哪个位置重启。作者定义了 expansion utility 指标,在 9 个模型 × 6 个基准共 41 个组合上逐步测量。结果显示重启位置确实影响正确率,5、16、38 个单元格的 held-out 审计中,基于位置的选择比均匀放置最多 +4.25 分。一个始终从最后一步重启的固定规则 always-last 表现很强,在 DeepSeek-R1-Distill-Qwen-14B / MATH-500 上以 0.774 倍的生成量超过四样本 self-consistency 上限 +0.052。
Expanding LLM Reasoning
Extra inference compute is usually spent on sampling more reasoning chains. We study where inside an existing chain an additional continuation should begin. We define expansion utility, the change in correctness from restarting a chain at a stored step, and measure it at every eligible step for nine models on six benchmarks (41 model and benchmark cells). Restart position matters: steps selected on one set of continuations beat uniform placement when scored on disjoint ones, in held-out audits on 5, 16, and 38 cells (+4.25 points [+2.51, +6.63] in a fresh five-cell audit). A fixed rule that restarts from the last eligible steps, always-last, is a strong baseline: our learned router beats uniform placement but shows no detected gain over it, and on DeepSeek-R1-Distill-Qwen-14B/MATH-500 always-last exceeds the exact self-consistency frontier at matched aggregate generated output by +0.052 [+0.008, +0.098], using 0.774x the aggregate generated output of four-sample self-consistency. Cross-fitted oracle selection still finds held-out headroom beyond declared positional classes, a target for future selectors. Finally, breaking step-label ties by earliest index flips the sign of a pointwise selector's gain over uniform placement in every seed of a five-seed diagnostic with four rollouts per step; randomized ties remove the bias.