A classifier's confidence score can tell the same story while its actual accuracy quietly falls apart.
That is the finding of a new paper examining how much a classifier's real-world accuracy can drift even when the distribution of its confidence scores never changes. The researchers model this using covariate shift - changes in the incoming data that don't touch the confidence outputs themselves - and define a "fragility profile" that measures the worst-case gap between what a model's confidence implies and what actually happens. Running the test on ImageNet classifiers, they found a measurable fragility gap in four of six primary classifiers and nine of twelve additional ones tested as released. After applying temperature scaling, a standard calibration fix, that number dropped to three of eighteen.
That gap matters because a confidence score is often treated as a stand-in for "trust this output." Here, the same confidence number stays tied to real-world accuracy that can swing hard under shifts in the input data, and the test shows temperature scaling isn't just cosmetic rescaling - it meaningfully shrinks that risk. For anyone shipping classifiers into production and treating calibration as a one-time checkbox, this is a reminder that a calibrated model on day one isn't guaranteed to stay honest as the data changes.
This is one unreviewed preprint measuring worst-case theoretical drift, not a documented failure in a live system - treat the fragility numbers as a stress test, not a prediction of what happens to your model next week.