AI/ ai research · interpretability · hallucination detection · retrieval-augmented generation

Researchers Find Where LLMs Signal Their Own Correctness

A new study locates a signal inside frozen language models that predicts whether their own answers are correct, no extra generation required.

Large language models apparently know when they are about to give you a wrong answer, and the tell is hiding in a single forward pass.

A new arXiv paper searches for exactly where that signal lives inside a frozen model, without generating extra text or sampling multiple answers. The researchers tested five families of internal readouts - hidden states, token probabilities, residual-stream features, attention, and combinations of all four - across two model backbones, in both closed-book question answering and retrieval-augmented settings. They found correctness concentrates in the answer span itself, recoverable from the answer tokens even when the model is pulling from retrieved documents. Combining signal types helped most when the test data looked different from training data, which is precisely when a single detector tends to fail.

Most hallucination detectors bolted onto retrieval-augmented systems check whether an answer matches the retrieved passage, not whether the answer is actually correct - a gap that shows up whenever the retrieved evidence is bad or contradictory. This work targets correctness directly, and cheaply: one forward pass, no extra generations, no separate classifier training loop. The authors also plug it into a retrieval controller, hinting at a use case beyond flagging bad answers after the fact - deciding when to bother retrieving at all.

It's a research result on two backbones, not a shipped feature, and every interpretability probe that looks tidy on a benchmark has a way of getting messier the moment it meets production traffic.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →