Turns out you can read a model's mind by checking its math homework.
Researchers studied vision-language models to see how much information survives as the model compresses its internal residual stream down to smaller representations. They compared two bottlenecks: a tuned-lens projection of the residual stream, and the model's final top-k logits, the handful of tokens most likely to show up in its answer. Both were tested on the same task, an image-based query where the model's actual answer is supposed to reveal only the requested information. The result: even the cheap, easily accessible top logits could leak details about the image that had nothing to do with the question asked, in some cases revealing nearly as much as inspecting the full internal representation directly.
Model builders often assume that what a user sees in the output marks the ceiling of what they can learn. This paper says that ceiling is porous even without fancy interpretability tooling. Just watching which tokens almost got picked can expose task-irrelevant information about an image, which matters for any product that treats a model's silence as a privacy or confidentiality guarantee.
If a handful of extra logits can spill what you tried to keep hidden, redacting the final answer was never the real safeguard.