AI/ ai research · multi-image ai · vision-language models · lvlm

A Fix for AI Models That Blend Multiple Images Together

Researchers built a training-free patch called FOCUS that stops vision-language models from confusing details across multiple images without retraining.

AI models that ace single-image questions often get confused once you show them two pictures at once.

A new paper describes a problem researchers call cross-image information leakage: when large vision-language models process several images together, visual details from one photo bleed into the model's read on another, and accuracy drops. The proposed fix, called FOCUS, is training-free and works on any existing model architecture: it masks all but one image with random noise, runs the model on that partially-masked input, and repeats the process so each image gets a turn as the only visible one. The resulting outputs are then combined and adjusted against a noise-only baseline to cancel out the cross-talk. Across several multi-image benchmarks, and even video understanding tasks, FOCUS reportedly boosted accuracy without any extra training.

Multi-image and video reasoning - comparing product photos, reading a sequence of charts, following a video clip - is where a lot of real-world AI products are headed, and it's also where today's models quietly fall apart. A fix that bolts onto an existing model with no retraining is cheap to deploy, which matters more than raw benchmark gains for teams already running these models in production.

Still, a noise-masking patch treats a symptom, not the disease: if attention layers default to blending images instead of keeping them separate, that's a structural quirk of how these models were trained, and FOCUS papers over it rather than resolving it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →