[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-fixing-ais-habit-of-overexplaining-easy-answers":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},9074,"fixing-ais-habit-of-overexplaining-easy-answers","Fixing AI's Habit of Overexplaining Easy Answers","A new self-distillation trick trims AI reasoning models' habit of writing long answers to questions they've already solved, without hurting accuracy.","A new training trick stops AI reasoning models from padding out answers to questions they already know how to solve.\n\nThe issue starts with reinforcement learning post-training, the fine-tuning stage where models practice solving problems and get rewarded for correct answers. Longer responses are often treated as a proxy for better reasoning on hard problems, but researchers found that same length creep leaks into easy problems too, with no accuracy benefit. They call this the length-scaling tax, or LST: extra words on already-solved queries that don't make the answer any better. Their fix, called Length Self-Distillation, routes easy prompts through a distillation step instead of standard reinforcement learning, while hard, unsolved prompts still get the normal training process. The twist is that the \"teacher\" guiding the distillation isn't a separate model at all, it's just a running average of the model's own earlier weights.\n\nThis matters because verbosity has a real price tag. Every extra sentence a model generates costs more compute, more latency, and more money at scale, and that cost compounds across millions of queries. A method that trims padding without needing extra data or an outside model is cheap to bolt onto training pipelines that already exist, which is probably why the gains are notable: the tax dropped from 19.0% to -3.7% on single-turn reasoning tasks, and from 31.4% to 13.7% on multi-turn agentic tasks, with accuracy holding steady or improving.\n\nStill, this is a patch on a training method, not a rethink of why models learned to pad in the first place, since the underlying incentive, rewarding length as a stand-in for effort, hasn't gone away.","[\"ai\",\"reasoning-models\",\"reinforcement-learning\",\"ai-efficiency\"]","2026-10-01T04:00:00.000Z","2026-10-01T18:34:03.088Z","2026-10-01T18:34:08.224Z","published",null,[],"ai",[24,26,27,28],"reasoning-models","reinforcement-learning","ai-efficiency",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38854",0,{"sections":35},[36,39,43,48,53,58,62,67,72,76,81,86,91,96],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",5488,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",809,{"name":44,"slug":45,"count":46,"latest_published_at":47},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":57},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",162,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":73,"slug":74,"count":70,"latest_published_at":75},"Software","software","2026-09-30T21:41:11.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]