AI/ ai · llm-reasoning · research · benchmarks

Preprint Questions Whether AI Math Gains Are Real Progress

An unreviewed arXiv preprint (2609.37066) finds AI math gains are sometimes real, sometimes just cheaper sampling of answers models already had.

AI models keep getting better at math competitions, but a new preprint asks whether that improvement is real reasoning or just cheaper luck.

The paper, arXiv:2609.37066 ("Beyond Compression: Diagnosing How Post-Training Changes Mathematical Reasoning"), is an unreviewed preprint posted September 30 with no named authors or institution listed. It compares three post-training methods for math-solving language models: the authors' own off-policy distillation, Alibaba's released Qwen3 distillation endpoints, and a DeepSeek-Math endpoint trained with Group Relative Policy Optimisation (GRPO), a reinforcement-learning technique. Instead of just checking whether a model gets the right answer once, the researchers ran each model many times per problem and tested it against paraphrased, translated, and numerically altered versions of the same questions, not just the exact wording it may have memorized.

The results split in two. On easier AMC-level problems, models were already close to solving everything possible, so more training just made correct answers cheaper to find, not more reachable. On harder AIME problems, training expanded what was actually solvable: Qwen3's endpoints pushed that ceiling highest, and DeepSeek's GRPO approach did not beat simple distillation at scale. English-heavy training also improved other-language performance without closing the gap between languages.

That distinction matters because most leaderboard pass@1 scores cannot tell smarter from luckier, which makes it hard to know whether a lab's post-training method teaches new reasoning or just repackages what the base model already knew.

Take it as a hypothesis worth testing, not a verdict. It is one uncredited preprint, not peer-reviewed work from a named lab.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →