AI/ vision-language-models · adversarial-attacks · ai-security · interpretability

Adversarial Attacks Fool AI Training, Not Its Answers

A new study finds adversarial attacks on vision-language models fail silently because the language decoder, not the image encoder, decides if an attack sticks.

A trick designed to make an AI vision model lie through its teeth doesn't actually work, and researchers now have a precise explanation why.

Researchers targeted Qwen2.5-VL-7B-Instruct, a vision-language model, with a crafted pixel-level perturbation meant to make the model internally accept a false caption during training. Using a two-stage PGD attack on 200 held-out COCO images, they drove the model's training-style loss for the fake caption to near zero. But when the same model was asked to freely describe the image, rather than being walked through the fake caption step by step, it described the image correctly every time, with no trace of the attack. Using a technique called the logit lens, which inspects a model's internal word rankings layer by layer, the team found the corrupted answer always sat at exactly the same rank, 3,488th out of 152,064 possible words, no matter which image was used.

Tracing that corruption through all 28 of the model's decoder layers, the researchers found the image-processing component damages every image's internal representation by roughly the same amount, whether the attack ultimately succeeds or not. The outcome gets decided later, inside the language model itself, which either amplifies the corrupted signal or actively suppresses it below its normal baseline. That means the model's language component, not its vision component, is effectively doing the defending.

It is a tidy reminder that a model's internal scratch work and its spoken answer can disagree, and that testing only the final output risks missing attacks that are working quietly under the hood, just not yet winning.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →