Researchers have built an attack that can blind the vision-language AI models used to read video feeds in self-driving systems.
The method, called Spatial Temporal Coherence Adversarial Attack (STCA), corrupts video before it ever reaches a vision-language model, and it works without any access to that model's internals. It runs in three stages: a caption-guided step picks out the most meaningful frames, a spatial-attack step adds perturbations crafted to stay visually close to the original footage, and a final stage uses a motion-guided mask to scramble how the model tracks movement across frames. The researchers tested it on the BDD100K and nuScenes driving datasets against three video-capable models, Video LLaVA-7B, Qwen2.5-VL-7B, and Dolphin. They trained the attack on one model and transferred it, unmodified, to fool the others.
The headline result is a high attack success rate while keeping the altered footage nearly indistinguishable from the original, measured by SSIM. That combination matters because a black-box attack doesn't require inside access to a target system, just a decent stand-in model to practice on, which is a much lower bar for anyone trying to interfere with a car's perception of the road.
It's a benchmark study on recorded video, not a camera mounted on a moving vehicle, and the paper's own conclusion is blunt: there's no real defense for this yet, just a warning that one is overdue.