AI models that ace single-image questions often get confused once you show them two pictures at once.
A new paper describes a problem researchers call cross-image information leakage: when large vision-language models process several images together, visual details from one photo bleed into the model's read on another, and accuracy drops. The proposed fix, called FOCUS, is training-free and works on any existing model architecture: it masks all but one image with random noise, runs the model on that partially-masked input, and repeats the process so each image gets a turn as the only visible one. The resulting outputs are then combined and adjusted against a noise-only baseline to cancel out the cross-talk. Across several multi-image benchmarks, and even video understanding tasks, FOCUS reportedly boosted accuracy without any extra training.
Multi-image and video reasoning - comparing product photos, reading a sequence of charts, following a video clip - is where a lot of real-world AI products are headed, and it's also where today's models quietly fall apart. A fix that bolts onto an existing model with no retraining is cheap to deploy, which matters more than raw benchmark gains for teams already running these models in production.
Still, a noise-masking patch treats a symptom, not the disease: if attention layers default to blending images instead of keeping them separate, that's a structural quirk of how these models were trained, and FOCUS papers over it rather than resolving it.