A popular way to catch AI models lying might not be catching lies at all.
Researchers tested linear probes, simple classifiers trained on a language model's internal activations, that claim to decode whether a model "knows" the truth even when it answers falsely out loud. They ran a game where a model has to report a secret bit honestly to an ally and lie about it to a rival. On ally-context training data, the true bit and the officially correct answer turn out to be the exact same label, so a probe trained there cannot tell which one it is actually tracking, a flaw the authors call "perfect aliasing." For a Gemma-2-9B model reinforcement-trained to lie on every rival trial, an ally-fitted probe scored just 0.006 AUROC on rival trials, which reads like proof the model buried the truth. But a probe trained on both ally and rival examples together scored a perfect 1.000 on those same activations, meaning the true bit was recoverable the whole time. A matching split showed up in Llama-3.1-8B, where two probes with identical 100 percent accuracy on ally prompts scored 0.080 and 0.986 AUROC on rival prompts, purely depending on how they were fitted.
Linear probes are a go-to tool in AI safety research for catching models that say one thing while representing another internally, and this paper shows the method can manufacture false evidence of deception, or false evidence that deception failed, based on how the probe's training labels happen to line up, not on anything the model actually did. That matters for anyone treating probe scores as a safety metric rather than a result that needs cross-checking. The authors also flag that their cleanest result used a game where the secret bit was spelled out in the prompt, so the probe may just be reading back a stored copy of it, a caveat that gets more serious once a model has to infer the bit itself.
It is a useful reminder that an interpretability tool measuring "truth" can end up measuring its own training setup instead, the machine-learning equivalent of a polygraph that flags nerves rather than lies.