A new attack shows that multimodal AI search tools can be tricked into leaking the private images they were built to protect.
Researchers built an attack called imMRAG against multimodal retrieval-augmented generation systems that return images directly as their answer, a setup used in medical-assistant and document-lookup tools. Instead of typing a malicious prompt, the attacker blends a shadow image with each image already pulled from the system, then uses the results to steer the next query toward unexplored parts of the embedding space. Testing against CLIP-family retrievers in medical, document, and general-purpose settings, a single run of 2,500 queries recovered up to 611 distinct radiology images, 566 document scans, and 416 general images. That is up to 5.6 times more unique items than a non-adaptive version of the same attack.
Retrieval-augmented generation is supposed to make AI answers more grounded and auditable by pulling from a known datastore instead of the model's memory. This research shows that same datastore becomes a target: an attacker never sees the data directly, never touches the model's weights, and still walks out with hundreds of images, including ones as sensitive as radiology scans.
Text-prompt jailbreaks have dominated the AI-safety conversation for two years; this is a reminder that the picture itself can be the attack vector, and most multimodal systems were not built to notice.