A new study on human-AI collaboration finds that AI models usually have a decent internal sense of when they're about to get something wrong - they just rarely tell you.
Researchers built a puzzle task where a human Helper tells an AI Worker where to place pieces, then tested three frontier vision-language models (GPT-4.1, GPT-5, and GPT-5.5). When asked to report a full probability distribution over candidate pieces, that signal was reasonably well-calibrated (an ECE, or calibration error, of 0.15, meaning the model's stated confidence closely tracked how often it was actually right) and useful for spotting mistakes (an AUROC of 0.65, a measure of how well a signal separates correct from incorrect placements, where 0.5 is a coin flip). The model's own token-level confidence, by contrast, was almost always near-certain (97% average) and poorly calibrated (ECE of 0.44). Despite having the better signal available internally, the models asked for clarification on only 3.5% to 16.7% of turns.
The stakes show up in a 210-person human study. Given only a model's default, unhedged message, people accepted 78% of wrong placements and could not tell good instructions from bad ones (an AUC of 0.50, meaning they were essentially guessing). Precise wording, and especially well-targeted hedges like "I think this piece, but I'm not fully sure," cut wrong-move acceptance to 36% without scaring people off correct moves. But a hedge generated automatically from the model's own uncertainty score did not reliably help, since it inherited the weaknesses of that underlying signal - sometimes making people trust bad advice more, not less.
That is the uncomfortable finding for anyone building AI copilots: an AI that "knows" it is unsure is not the same as one that tells you usefully. A hedge is only as good as its calibration, and a poorly calibrated hedge is worse than silence.