Ask a chatbot how sure it is, and the answer may not be pure improvisation: new research finds a model's stated confidence actually tracks its internal math.
Researchers tested two separate signals in large language models: the internal probability distribution a model uses to pick one answer over another, and the confidence it states out loud when asked directly. They deliberately manipulated two sources of uncertainty in training and in-context data, including how often an answer appears (distributional frequency) and explicit statements asserting a probability. Both signals shifted in response to each type of manipulation. Crucially, the internal and verbalized numbers moved together more closely than expected if each were simply tracking the same training data independently.
That coupling is the real finding. It means a model's spoken confidence level is not just a plausible-sounding phrase; it is a usable, if imperfect, proxy for the probability distribution the model is actually running internally, something researchers previously could not assume held true. Anyone building systems that lean on self-reported confidence, flagging shaky answers, routing uncertain cases to a human, calibrating automated decisions, now has evidence that number is worth measuring rather than discarding.
One caveat: coupled is not identical. Models can still state a confident number while being flatly wrong, so treat verbalized uncertainty as a useful hint, not a guarantee.