A new arXiv paper argues the best judge of an AI model's hallucinations might be a different AI model entirely.
Researchers built a framework that inspects a language model's layer-wise internal activations to flag not just that a hallucination happened, but exactly which tokens started it and how far it spread. That is a step up from most existing internal-state probes, which treat hallucination detection as a crude yes-or-no call on each token. The team then tested a cross-model setup: one model watches the hidden states produced while a second model generates text, rather than checking its own output. The observer model matched or beat the generator's own self-detection of hallucination onsets, and this held even when the observer was the smaller of the two models.
That last point is the real finding. It implies self-monitoring is not actually the ceiling for catching hallucinations early, and that a cheap, bolted-on watchdog model could audit a bigger, more expensive one without needing external fact-checking lookups. For companies trying to make LLM output trustworthy without the latency and cost of retrieval-based verification, that is a meaningfully different architecture to consider.
It is one preprint with results measured against benchmark precision-recall curves, not a deployed safety system, and beating a random baseline under heavy class imbalance is a lower bar than reliably catching hallucinations across messy, open-ended prompts in production.