AI vision models can miss half the words hiding in a single image, and a new benchmark just proved it with two fonts stacked on top of each other.
Researchers built a 300-image dataset called DecoyBench using a method they call Decoy Font: each image pairs sharp, high-contrast text with a second layer of soft, blurred text sitting underneath it. Six closed-source vision-language models (VLMs) from three different model families were shown the images under two prompting setups - one that just asked what the image said, and one that explicitly told the model two layers were present - at both a normal resolution (512x512 pixels) and a tiny one (64x64 pixels). Human volunteers who looked at the same images read both layers of text with high accuracy, no matter the setup.
The models did not. At high resolution, most of them read the sharp-edged text about as well as humans do, but almost never pulled out the soft, blurred text underneath, even when told it was there. Shrink the image down and the pattern flips: the sharp text becomes unreadable to humans and machines alike, while the models suddenly read the blurred text just fine.
That flip is the real finding. It is not that these models are simply bad at reading blurry or overlapping text - they consistently favor one kind of visual detail over another, in a way humans do not. That is a structural blind spot baked into how the models process an image, not a resolution problem, and it shows up the same way across model families and prompting styles.
Any product that leans on a VLM to read what is actually in an image - screenshots, scanned forms, content moderation queues - now has a documented, repeatable way to make it see only half the message.