[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-training-method-patches-a-flaw-in-ai-reasoning-models":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},9216,"new-training-method-patches-a-flaw-in-ai-reasoning-models","New Training Method Patches a Flaw in AI Reasoning Models","A new reinforcement-learning tweak that filters out irrelevant wording cues lifts a model's math accuracy by several percentage points across benchmarks.","A new training tweak quietly fixes a bug that's been inflating how smart AI reasoning models look on math tests.\n\nResearchers tested large language models trained with reinforcement learning on verifiable rewards - the method behind today's best reasoning models - by making cosmetic changes to math problems that left the underlying problem and answer untouched. They found certain words shift wildly in probability when the wording changes, even though nothing about the actual problem did, and that this \"token drift\" quietly degrades reasoning. The culprit is Group Relative Policy Optimization (GRPO), the standard training algorithm, which hands out the same reward or penalty to every token in a response, including the ones a model latched onto for no good reason. The team's fix, called Semifactual Credit-Augmented Policy Optimization, tracks how much each token's probability drifts under those cosmetic prompt changes and reduces the training credit given to unstable tokens early on.\n\nOn two Qwen3 models, the new method beat standard GRPO training by 5.63 and 4.17 percentage points on the AIME math competition benchmark, and it won on most other math tests and every out-of-distribution test the team tried. That's a useful reminder that some of what looks like \"reasoning\" in these models is pattern-matching on phrasing, and that stripping out that brittleness is worth real accuracy points, not just a theoretical concern.\n\nIt's plumbing work on an existing training recipe, not a new model - but plumbing like this has a habit of showing up quietly in next year's leaderboard numbers.","[\"reinforcement-learning\",\"llm-reasoning\",\"ai-training\",\"benchmarks\"]","2026-10-01T04:00:00.000Z","2026-10-02T02:22:22.746Z","2026-10-02T02:22:28.086Z","published",null,[],"ai",[26,27,28,29],"reinforcement-learning","llm-reasoning","ai-training","benchmarks",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.40360",0,{"sections":36},[37,40,44,48,53,58,62,67,72,76,81,86,91,96],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",5612,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",815,{"name":45,"slug":46,"count":47,"latest_published_at":18},"Policy","policy",430,{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":57},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",163,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":73,"slug":74,"count":70,"latest_published_at":75},"Software","software","2026-09-30T21:41:11.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]