探索广度与推理精度:通过采样提升小模型性能
Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
新PPT方法让小模型通过并行采样达到前沿性能,比强化学习训练更高效。
研究人员提出并行功率退火(PPT)方法,通过在不同锐化级别并行运行多个交互副本,解决探索-利用权衡问题。该方法在推理时提升大语言模型推理能力,无需参数更新或外部奖励。实验显示PPT显著提升单链功率锐化采样性能,甚至达到前沿模型水平。
Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or external rewards, avoiding the costly optimization and jagged generalization of RL. However, this approach faces a fundamental exploration--exploitation trade-off, as % strong sharpening restricts exploration, trapping samplers in plausible but incorrect reasoning trajectories, whereas weak sharpening leaves the answer distribution diffuse. To resolve this trade-off, we introduce \textbf{Parallel Power Tempering (PPT)}, instantiating power-sharpened LLM sampling via parallel tempering. Running multiple \emph{interacting} replicas in parallel at different sharpening levels allows lower-power replicas to explore diverse reasoning trajectories and higher-power chains to further exploit higher-likelihood responses favored by the sharpened target. Specifically, we tailor \method{} to inference-time sampling by mitigating a truncation bias, identified in prior power samplers, and investigate effective swap strategies under finite memory and compute budgets. Extensive experimentation shows that \method{} substantially improves single-chain power-sharpened sampling and outperforms RL-post-trained models, producing higher-quality reasoning traces and even achieving performance comparable to frontier models.