AI/ llms · math-reasoning · benchmarks · ai-research

LLMs Still Struggle When Math Problems Get Reframed

A new benchmark rewriting Putnam problems finds top AI models fail when a proof's setting changes, not just its wording.

A new benchmark shows top AI models can still be tripped up by math problems that only look different.

Researchers built a system called GAP that takes existing math problems and generates equivalent variants in two ways: renaming the labels in a problem, or changing the underlying mathematical setting while keeping the same proof idea. They applied this to all 1,051 Putnam Competition problems from 1938 to 2024, adding 5,255 new variants to create a 6,306-item dataset called PutnamGAP. They then tested 18 commercial and open-source models against it. Accuracy fell across every model, and the drop was worst when the mathematical setting changed rather than just the names.

The result suggests models that score well on existing math benchmarks may be recognizing patterns from training data rather than genuinely reasoning through a proof. The gap did not shrink for the strongest models, meaning bigger or newer isn't fixing the real weakness: carrying a proof strategy over to a new context.

Olympiad-level accuracy numbers look impressive in a press release, but a correct answer and a general method are not the same thing.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →