Security/ ai · security · jailbreak · multimodal-models

New Attack Hides Harmful Prompts Inside Video Scenes to Jailbreak AI

Researchers found that wrapping a harmful question in an innocuous video story fools GPT-4.1 and Gemini 3.5 Flash into answering it anyway.

A new jailbreak technique hides harmful requests inside ordinary-looking video scenes, and it works on some of the best multimodal AI models available.

Researchers built a framework called SceneJail that searches for a video scenario compatible with a harmful question, then uses feedback from the model's own answers to refine the text prompt around it. They tested the technique on eight video-capable multimodal models, including GPT-4.1 and Gemini 3.5 Flash, using two existing jailbreak benchmarks. One version, which presents the full harmful query persistently through the video, hit an average attack success rate of 91.5 percent, beating the strongest prior method by 29.1 percentage points. A second version spreads the query across successive frames instead, and it kept a 72.3 percent success rate even against strict image filtering meant to catch this kind of attack.

Most video jailbreak work until now has focused on hiding harmful content within a single frame or image, treating the video as just a delivery mechanism. SceneJail instead treats the surrounding story, the scenario the video builds around the request, as its own attack surface, and that is harder for a filter to catch because nothing in any single frame looks obviously harmful.

Image-level filters have gotten decent at spotting what is bad in a frame. Nobody has really built a filter for a story that is setting you up, and this paper is a reminder that one is overdue.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →