AI/ hallucination · large language models · ai research · nlp

Study Finds LLMs Know When They Are Guessing But Guess Anyway

New research shows large language models can detect their own uncertainty yet still choose specific, invented details over safer, vaguer answers.

Researchers have shown that large language models secretly know when they are about to make something up, and generate the fabrication anyway.

Using a benchmark built on the T-REx knowledge base, researchers tested whether LLMs can tell when an entity falls outside what they actually know, and whether they can anticipate how specific an answer is about to be. Probing the models' internal activations, they found both signals are present: models do encode knowledge-boundary awareness and do anticipate referent specificity before generating. But the models never act on that awareness. Faced with an unfamiliar entity, they still favored a specific-sounding answer over a vaguer, safer one, even when a correct generic alternative was available to them.

That reframes hallucination as less of a knowledge problem and more of a policy problem, a mismatch between what a model represents internally and what it actually outputs. A cooperative human speaker unsure of a detail retreats to something vaguer but true, like "some kind of manager" instead of guessing a title. LLMs have the internal ingredients for that retreat and skip it anyway, defaulting to specificity over honesty.

That is a more precise diagnosis than the usual hand-wave about models hallucinating, though turning a diagnosis into a fix, what the paper calls Gricean alignment, is still a proposed training objective, not a shipped model.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →