[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-training-method-teaches-models-when-to-stop-thinking":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},9260,"ai-training-method-teaches-models-when-to-stop-thinking","AI Training Method Teaches Models When to Stop Thinking","A reinforcement-learning technique called TRAAC cuts reasoning length by over a third while boosting accuracy on math and science benchmarks.","A new training method called TRAAC teaches AI reasoning models to think less on easy problems and more on hard ones.\n\nReasoning models solve tough problems by scaling up test-time compute, essentially thinking longer before answering. The catch: models often get this wrong in both directions. Short reasoning causes errors on hard problems, while needlessly long reasoning wastes tokens restating a correct answer the model already found. TRAAC addresses this with reinforcement learning applied after initial training. It uses the model's own attention patterns to spot which reasoning steps matter and prune the rest, and it factors in an estimate of problem difficulty so the model learns to size its reasoning budget to the task. Tested on the Qwen3-4B model across AIME, AMC, GPQA-Diamond, and BBEH benchmarks, TRAAC delivered an average 8.4 percentage point accuracy gain and a 36.8 percent cut in reasoning length compared to the base model, with similar gains holding on benchmarks it wasn't trained for.\n\nThe real cost of reasoning models isn't just compute, it's trust. Users have learned to be skeptical of \"thinking\" animations that grind on with no payoff, and every extra token of unnecessary reasoning is money spent for nothing. TRAAC's trick, compressing redundant steps instead of just capping token counts, points to a way of making inference cheaper without the usual accuracy tradeoff that comes with simply forcing shorter answers.\n\nThe gains are real but come from a single 4 billion parameter model family tested on a handful of benchmarks. Whether the same attention-based pruning holds up on the much larger reasoning models already shipping in production is still untested.","[\"ai\",\"reasoning-models\",\"machine-learning\",\"research\"]","2026-10-01T04:00:00.000Z","2026-10-02T05:21:26.106Z","2026-10-02T05:21:28.965Z","published",null,[],"ai",[24,26,27,28],"reasoning-models","machine-learning","research",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2510.01581",0,{"sections":35},[36,39,43,47,52,57,61,66,71,75,80,85,90,95],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",5671,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",820,{"name":44,"slug":45,"count":46,"latest_published_at":18},"Policy","policy",430,{"name":48,"slug":49,"count":50,"latest_published_at":51},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":18},"Science","science",165,{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":72,"slug":73,"count":69,"latest_published_at":74},"Software","software","2026-09-30T21:41:11.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]