[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-training-cuts-ai-overclaiming-from-97-to-35-not-to-zero":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},10043,"training-cuts-ai-overclaiming-from-97-to-35-not-to-zero","Training Cuts AI Overclaiming From 97% to 35%, Not to Zero","Fine-tuning and reinforcement learning cut an AI investigator's tendency to overstate conclusions from 97% to 35%, but the problem didn't disappear.","A new training method cuts how often AI incident investigators declare a case closed without enough evidence - but it does not come close to solving the problem.\n\nResearchers built Nautil, a set of 731 audited investigation reports covering aviation, rail, maritime, chemical-safety and vehicle-defect incidents plus production server outages, complete with teacher examples of good reasoning. Before training, an off-the-shelf 9-billion-parameter model overstated its conclusions in 97% of answers. A frontier model fared little better: it named the right cause 84% of the time, still overstated 91% of the time, and closed 17 of 41 cases that official investigators had labeled cause undetermined. After fine-tuning the 9B model on Nautil's trajectories, overstatement dropped to 35% and correct, properly hedged conclusions rose from 3% to 43%. A follow-up reinforcement-learning step, rewarding only the close-or-keep-open decision, pushed balanced accuracy from 69.2 to 83.3, matching the quality of the human-curated training examples.\n\nWhy it matters: incident reports get treated as authoritative, and a model that confidently names a cause when the evidence does not support it could slot into real safety workflows without anyone catching the bluff. This work targets a quieter failure mode than most AI safety research - not hallucinated facts, but premature certainty - and the baseline numbers show how badly models default to overconfidence without specific training against it.\n\nCutting the overstatement rate from 97% to 35% is a real improvement, not a cure. An AI investigator that is still wrong more than a third of the time is not ready to close any case on its own.","[\"ai-safety\",\"llm-evaluation\",\"reinforcement-learning\",\"ai-research\"]","2026-10-05T04:00:00.000Z","2026-10-05T20:02:30.375Z","2026-10-05T20:02:36.368Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The headline claims the training method 'Stops' AI from overclaiming, but the body reports overstatement only fell from 97% to 35% (still common) — rewrite the headline\u002Fdek to reflect a reduction, not an elimination, of the problem.","resolved","ai",[32,33,34,35],"ai-safety","llm-evaluation","reinforcement-learning","ai-research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.03190",0,{"sections":42},[43,46,50,55,60,65,69,74,78,83,88,93,98,103],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",6290,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",869,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",444,"2026-10-03T15:02:01.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",323,"2026-10-04T13:00:00.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",204,"2026-10-03T14:50:50.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":18},"Science","science",178,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",158,"2026-10-03T03:21:12.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":18},"Dev Tools","dev-tools",98,{"name":79,"slug":80,"count":81,"latest_published_at":82},"Software","software",97,"2026-10-04T10:00:00.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",92,"2026-10-04T14:36:25.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"General","general",51,"2026-10-05T02:35:01.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",32,"2026-10-02T18:00:00.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]