A new study says the confidence scores AI systems report about their own answers don't reliably update as the underlying model's knowledge changes.
Researchers studied what they call "persistent calibration": whether a confidence estimator trained on one checkpoint of an open model still works on a later checkpoint of the same model. To test it, they built "knowledge contrast sets" - questions one checkpoint answers correctly and another answers incorrectly - and checked whether confidence scores actually tracked those shifts. Both inference-time methods and fine-tuned estimators fell short on these contrast sets, even when they looked well-calibrated across the full dataset. Only estimators trained with access to future checkpoints performed reliably, though training on multiple checkpoints at once narrowed the gap.
This matters because AI products are increasingly pitched as agents that keep learning and updating, not static tools. A confidence score is only useful if it reflects what a model currently knows, not what it knew when the score-generating method was built. The paper's evidence suggests most calibration techniques are tuned to a specific model snapshot rather than general-purpose self-knowledge.
So the next time a model says it's 90 percent sure, treat that as a snapshot, not a subscription - it may need retraining every time the model underneath it changes.