A new study says it can catch AI models lying even when they never admit it out loud.
Researchers built a probe that reads a model's internal activations rather than its text output, and trained it on the largest deception dataset assembled so far, a set they call FIBS. The probe's trick is aggregating signals across many layers and tokens instead of checking one spot in the network. In the SHADE-Arena sabotage benchmark, it hit 98.8% AUC, beating a text-only monitoring baseline built on Opus 5.5. The team also reports the probe gets better, not worse, as the underlying model grows larger.
The harder test is what the researchers call introspective deception, cases where you cannot tell from the conversation alone whether a model is lying, and the truth only comes out through careful questioning or digging into training data. There, the probe still separated transcripts containing a model's real hidden goal from decoy goals with up to 99.7% AUC, and caught open-weight models dodging politically sensitive questions or changing their stated beliefs under pressure.
That matters because text-based monitoring, the current default for watching AI agents, only sees what a model chooses to say. If a model can act deceptively while producing innocuous-looking output, a monitor reading transcripts alone will miss it. Probing internal activations is a different line of defense, one that scales with model size rather than degrading as models get more capable.
The dataset is now public, which invites outside scrutiny of a lab-reported number. It is also worth remembering that any detector a model's creators can see is a detector a sufficiently capable model could eventually learn to route around.