Study finds large language models often miss the core math idea that makes a solution work
Researchers introduce a new way to probe whether large language models really understand mathematics or only follow steps that happen to lead to correct answers. They define a “Mathematical Primitive” as the short, problem-specific idea that explains why a solution works. Using this concept, they build a benchmark that tests four things separately: whether a model can discover the primitive from the problem, generate a full solution, extract the primitive from a provided solution, and execute a solution when given the primitive explicitly.
To make this benchmark, the team built on a curated set of hard math problems from a collection known as Humanity’s Last Exam. They used a strong model (GPT-5.4-High) to draft primitive annotations and then had three human experts with graduate-level math training review and verify each item. The benchmark groups primitives into families (called Recast, Witness, and Argument) and scores predicted primitives against human-annotated gold standards using a rubric that checks validity, whether the core structure was identified, and whether the explanation shows how that structure enables the solution.
Their diagnosis shows that final-answer accuracy can hide very different internal abilities. In many cases a model can carry out calculations or write a correct-looking derivation without ever identifying the compact idea that makes the problem solvable. Conversely, when the correct primitive is provided, models often show a large latent capacity to execute the remaining steps. From their experiments, the single biggest bottleneck is Discovery — the model’s ability to spot the primitive from the problem alone.
The authors also study post-training fixes. They find that failures caused mainly by missing discovery are more amenable to repair than other kinds of failures. Building on that, they introduce a post-training method that uses primitives as privileged information during training and selectively distills primitive-guided reasoning into a student model. This method does not require primitives at inference time and, across model sizes and challenging benchmarks, yields consistent improvements in average performance of about 2.42 to 4.48 points over baselines.