The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
A new benchmark shows LLMs lack structural mathematical understanding with discovery as the key bottleneck, and a primitive-guided self-distillation framework repairs reasoning to boost performance.
Published Oct 1, 2026▲ 10 on Hugging FaceCode ★ 8arXiv ↗

Only vote on papers you've read. Sign in with GitHub to vote.
A rigorous benchmark exposes hidden reasoning profiles and shows discovery failures are repairable, though whether primitives represent true structural understanding or just a sharper taxonomy remains debated.
Abstract
While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce \abs{}, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that \abs{} consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.