Security/ ai-safety · jailbreaks · vision-language-models · security-research

Researchers Find No Safe Middle Ground for VLM Jailbreak Defenses

A new study finds recovery-and-reguard defenses cut some VLM jailbreaks but no setup blocks most attacks without also flagging normal users.

A new study on vision-language model safety guards finds that patching one jailbreak hole just moves the attack somewhere else.

Guards that screen prompts for vision-language models normally judge an input's surface form, not its actual meaning - so a harmful request re-encoded as set theory, formal logic, a classical language, code, or text hidden inside an image can slip straight past them. Researchers built a "recover-and-reguard" pipeline that first recovers the image or decodes the encoding, then re-checks it, and tested it against eleven encoding attacks: six published implementations, one standard baseline, one adapted method, and three attacks they built themselves. Restoring what the guard never saw worked on its own terms - block rates on image-based renders jumped from zero to 67-90%. But the cost in false blocks on ordinary requests varied wildly: one guard's defense cost it 9 extra percentage points of over-refusal for a 70-point security gain, while another guard paid 69 points for the same gain.

Against an attacker free to pick any of the eleven encodings, though, closing one channel just relocates the jailbreak instead of closing it - the researchers found no ensemble-wide improvement in attack success that survived statistical correction. Adding a "reguard" step that re-screens the recovered content before decoding did lower overall attack success, but it was the one fix that raised over-refusal rates across every single guard-and-target pairing tested. Across the full matrix, no combination of guard, target, and defense method got attack success at or below 40% while keeping benign over-refusal under 70%.

The authors call that empty middle ground a property of the setups they tested, not a hard ceiling. Until someone proves otherwise, a "safer" VLM guardrail mostly seems to mean one that's worse at telling real users from attackers.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →