Auditing an AI model's internal activations for hidden threats doesn't require reading every token it produced.
The paper, posted to arXiv as arXiv:2609.37040 and not yet peer-reviewed, tests methods for scanning natural language autoencoders - tools that translate a model's internal activations into human-readable explanations, used to catch things like prompt injection attempts or concealment. Across 4.7 million explanations, the authors found a ranker based on chat structure alone picked out useful positions to explain better than signals pulled from the model's own computation, without needing to run the model at all. On three of the four datasets tested, explaining just 5% of token positions retained nearly all of the detection success rate achieved by explaining every position; the size of that benefit varied depending on the audit task. The team also found that off-the-shelf verbalizers - the components that turn activations into words - could surface terms a model had been fine-tuned to conceal, with no extra training required.
Explaining every token in every prompt is the kind of interpretability tax that keeps safety audits from scaling past a research lab's budget. If a cheap heuristic that skips the forward pass gets auditors most of the way to full coverage on most tasks, that shifts interpretability from a luxury line item toward something closer to routine monitoring. It also hints these tools generalize past the specific model a verbalizer was trained on, which matters if audit tooling built for one model ends up checking another.
One dataset didn't cooperate, and the paper's own numbers show the payoff swings by task. Treat "5% coverage, most of the results" as a promising shortcut worth watching, not a settled result - this hasn't been peer reviewed yet.