AI/ ai-safety · llm-evaluation · reinforcement-learning · ai-research

Training Cuts AI Overclaiming From 97% to 35%, Not to Zero

Fine-tuning and reinforcement learning cut an AI investigator's tendency to overstate conclusions from 97% to 35%, but the problem didn't disappear.

A new training method cuts how often AI incident investigators declare a case closed without enough evidence - but it does not come close to solving the problem.

Researchers built Nautil, a set of 731 audited investigation reports covering aviation, rail, maritime, chemical-safety and vehicle-defect incidents plus production server outages, complete with teacher examples of good reasoning. Before training, an off-the-shelf 9-billion-parameter model overstated its conclusions in 97% of answers. A frontier model fared little better: it named the right cause 84% of the time, still overstated 91% of the time, and closed 17 of 41 cases that official investigators had labeled cause undetermined. After fine-tuning the 9B model on Nautil's trajectories, overstatement dropped to 35% and correct, properly hedged conclusions rose from 3% to 43%. A follow-up reinforcement-learning step, rewarding only the close-or-keep-open decision, pushed balanced accuracy from 69.2 to 83.3, matching the quality of the human-curated training examples.

Why it matters: incident reports get treated as authoritative, and a model that confidently names a cause when the evidence does not support it could slot into real safety workflows without anyone catching the bluff. This work targets a quieter failure mode than most AI safety research - not hallucinated facts, but premature certainty - and the baseline numbers show how badly models default to overconfidence without specific training against it.

Cutting the overstatement rate from 97% to 35% is a real improvement, not a cure. An AI investigator that is still wrong more than a third of the time is not ready to close any case on its own.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →