Security/ ai · security · prompt-injection · llm-safety

Defense Turns Text Into Images To Stop Prompt Injection

A new technique called Pictionary renders untrusted text as images before feeding it to AI models, exploiting a blind spot to cut prompt injection attacks.

A new paper shows that turning malicious text into a picture makes AI models less likely to obey it.

Researchers tested ten multimodal large language models and found a consistent asymmetry: they are more likely to follow adversarial instructions embedded in text than the exact same instructions delivered as an image or audio clip. Building on that gap, the team built a training-free defense called Pictionary, which renders any untrusted content - like a webpage snippet or tool output - as a typographic image before it reaches the model. Across two prompt injection benchmarks, DirectInject and AgentDojo, Pictionary cut attack success rates, holding up even against adaptive attacks and human red teamers, while barely denting the models' usefulness on legitimate tasks.

The finding suggests today's models do not have a general sense of "instructions I should follow" - they have just been trained to treat text as command-shaped and other modalities as things to merely describe. That is a fixable training artifact, not a fundamental limit: the researchers show that fine-tuning models on image-rendered instructions erodes the gap, tracing the vulnerability back to text-centric instruction-tuning data.

It is a clever stopgap, not a cure. The defense works because current models are bad at reading images as commands - once labs train models to follow image-based instructions too, as they eventually will, this particular blind spot closes.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →