论文精选

后训练如何改变数学推理能力

Beyond Compression: Diagnosing How Post-Training Changes Mathematical Reasoning

精选理由

这篇论文分析了不同后训练方法对数学推理的影响,发现压缩只是后训练的一种模式,而非普遍解释。

研究比较了三种后训练方法在数学推理上的表现:离策略蒸馏、Qwen3端点和DeepSeek-Math GRPO。在较简单的AMC问题上,后训练主要压缩了样本成本;在较难的AIME问题上,后训练扩展了大K上限。离策略蒸馏已能提高上限,Qwen3进一步提升了这一上限,而DeepSeek-Math GRPO在大K时并不占优。

原文 · arXiv: DeepSeek

Beyond Compression: Diagnosing How Post-Training Changes Mathematical Reasoning

Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed. Gains may reflect newly reachable solutions, cheaper sampling of latent solutions, surface robustness, or memorisation. We compare three post-training paths under a common diagnostic readout: our sufficiently trained off-policy distillation trajectories, released Qwen3 off-policy-plus-on-policy distillation endpoints, and a released DeepSeek-Math endpoint trained with Group Relative Policy Optimisation (GRPO). Our probe uses cross-surface pass@K over verbatim prompts, paraphrases, numerical isomorphisms, and translations, plus consistency, distribution-shape, and verified supervised-fine-tuning (SFT) membership analyses. We find two regimes. On easier AMC problems, large-K ceilings are near saturation, so post-training mainly compresses sample cost. On harder AIME problems, post-training expands the large-K ceiling over the base model: sufficient off-policy distillation already raises this ceiling, Qwen3 released endpoints raise it further, and DeepSeek-Math GRPO does not dominate sufficient off-policy distillation at large K. English-dominant distillation improves non-English reasoning but preserves language-tier gaps. A controlled-overfit audit finds limited sensitivity in current SFT-membership probes. Compression is one regime of post-training, not a universal explanation.