论文

诊断与修复大语言模型的数学推理能力

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

精选理由

这篇论文揭示了LLM数学推理的瓶颈,提出了新评估框架和改进方法,对提升AI数学能力有重要参考价值。

研究人员提出数学原语(Mathematical Primitive)概念,系统评估大语言模型的数学理解能力。他们开发了hlei基准,从发现、生成、消化和执行四个维度评估数学推理。诊断显示,解决方案准确性掩盖了不同的能力特征,原语能释放大量执行能力,而发现能力是数学推理的主要瓶颈。基于这些发现,研究人员开发了abs框架,通过选择性转移原语引导的推理来改进模型。

原文 · arXiv cs.LG

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce \abs{}, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that \abs{} consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.