Hallucination detectors don't fail evenly - and a new arXiv paper puts a number on just how uneven.
Researchers tested sampling-based consistency detection, the common technique of asking a model the same question multiple times and checking whether its answers agree, across four language models and three factual question-answering datasets. They split hallucinations into two groups: "Ghost" cases, where the model's repeated answers mostly agree even though the answer is wrong, and "Flickering" cases, where answers disagree a lot. The gap between how detectable these two groups are came out to 0.35-0.46 AUC (area under the curve), a 0-to-1 score for how well a detector separates hallucinated answers from correct ones, where 0.5 is a coin flip and 1.0 is perfect. Because the math used to define the groups was itself correlated with the math used to measure the gap, the team re-ran the test after locking in the group assignments, and the same asymmetry showed up in independent measures of answer wording and meaning, holding up across all 12 model-and-dataset combinations and surviving a stricter check on two additional model families.
That means the same detector can look reliable on average while quietly missing an entire class of errors on a given model. The share of hallucinations landing in the harder-to-catch group ranged from 16% to 77% depending on which model was tested, and the same prompt often flipped between easy and hard depending on the model answering it.
Any vendor citing a single hallucination-detection score is telling you less than it sounds like - the interesting number is the one they're not showing you: how that score splits by model and by case.