AI/ ai · benchmarks · math-reasoning · llm-evaluation

New Benchmark Shows Gemini's Math Reasoning Collapses Under Pressure

A new math benchmark shows Gemini-3.1-pro-preview, its standard-mode leader, scores below random guessing once questions are reworded, unlike GPT-5.4.

A new benchmark for advanced math reasoning finds that leading AI models still struggle to tell real proofs from clever fakes.

Researchers built LiveMathematicianBench by pulling theorems from arXiv papers published after the models' training cutoffs, so none of the questions could have been memorized. Each question sorts into one of thirteen theorem types, such as existence or uniqueness proofs, and the wrong answers are generated from real but flawed proof strategies rather than random noise. The benchmark also runs a substitution-resistant version, which rewords problems to check whether a model is actually reasoning or just pattern-matching to familiar phrasing. On the standard test, Gemini-3.1-pro-preview leads with 43.5% accuracy - already unimpressive on a multiple-choice format.

The real story is what happens when the wording changes. Gemini-3.1-pro-preview, the standard-mode leader, collapses to 17.6% under substitution-resistant testing - worse than random guessing, which sits at 20% on this format. GPT-5.4 takes over as the substitution-resistant leader at 30.6%, still well above chance but far from proof.

Proof-sketch access still boosted every model's scores, but a benchmark where the top model on one test falls below a coin flip on another is a reminder that leaderboard rankings shift depending on how hard you make the questions to game.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →