[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-preprint-questions-whether-ai-math-gains-are-real-progress":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8617,"preprint-questions-whether-ai-math-gains-are-real-progress","Preprint Questions Whether AI Math Gains Are Real Progress","An unreviewed arXiv preprint (2609.37066) finds AI math gains are sometimes real, sometimes just cheaper sampling of answers models already had.","AI models keep getting better at math competitions, but a new preprint asks whether that improvement is real reasoning or just cheaper luck.\n\nThe paper, arXiv:2609.37066 (\"Beyond Compression: Diagnosing How Post-Training Changes Mathematical Reasoning\"), is an unreviewed preprint posted September 30 with no named authors or institution listed. It compares three post-training methods for math-solving language models: the authors' own off-policy distillation, Alibaba's released Qwen3 distillation endpoints, and a DeepSeek-Math endpoint trained with Group Relative Policy Optimisation (GRPO), a reinforcement-learning technique. Instead of just checking whether a model gets the right answer once, the researchers ran each model many times per problem and tested it against paraphrased, translated, and numerically altered versions of the same questions, not just the exact wording it may have memorized.\n\nThe results split in two. On easier AMC-level problems, models were already close to solving everything possible, so more training just made correct answers cheaper to find, not more reachable. On harder AIME problems, training expanded what was actually solvable: Qwen3's endpoints pushed that ceiling highest, and DeepSeek's GRPO approach did not beat simple distillation at scale. English-heavy training also improved other-language performance without closing the gap between languages.\n\nThat distinction matters because most leaderboard pass@1 scores cannot tell smarter from luckier, which makes it hard to know whether a lab's post-training method teaches new reasoning or just repackages what the base model already knew.\n\nTake it as a hypothesis worth testing, not a verdict. It is one uncredited preprint, not peer-reviewed work from a named lab.","[\"ai\",\"llm-reasoning\",\"research\",\"benchmarks\"]","2026-09-30T04:00:00.000Z","2026-09-30T15:13:28.403Z","2026-09-30T15:13:34.998Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Add verifiable attribution — cite the paper's arXiv identifier (2609.37066) and title, and note it's an unreviewed preprint with no named authors\u002Finstitution given, since the draft currently only says 'a new arXiv study' with nothing a reader could use to verify or look up the source.","resolved","ai",[30,32,33,34],"llm-reasoning","research","benchmarks",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.37066",0,{"sections":41},[42,45,49,53,58,63,68,73,78,82,87,92,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5135,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",788,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",417,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]