Researchers have a fix for AI models that sound certain when they shouldn't be: make the model's reasoning double-check its own confidence.
The method, called RL-ARC, targets a known side effect of reinforcement learning with verifiable rewards, or RLVR, a popular technique for training models to reason better by rewarding correct final answers. The problem is that RLVR does not account for calibration, meaning the gap between how confident a model sounds and how often it is actually right. Left unchecked, this training approach pushes models toward overconfidence. Prior fixes that bolt on uncertainty-estimation objectives help some, but researchers found they still produce overconfident answers when a model faces questions unlike its training data, and often cost some reasoning ability in the process. RL-ARC instead uses the model's confidence in its own reasoning steps as a check on its confidence in the final answer, rewarding calibrated confidence when the reasoning is correct and penalizing overconfidence when it is not.
This matters because calibration, not raw accuracy, is often what determines whether an AI system is safe to deploy. A model that is wrong 10 percent of the time is manageable if it flags that 10 percent with lower confidence; a model that is wrong 10 percent of the time while sounding equally sure every time is not. According to the researchers, RL-ARC improved calibration in both familiar and out-of-distribution settings without the substantial reasoning-performance tradeoff seen in earlier calibration-aware training methods.
It's a narrow fix for a problem that keeps resurfacing wherever reasoning models get pushed into real use: they are trained to be right, not to know when they might be wrong, and those turn out to be different skills to teach.