[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-vision-models-miss-hidden-text-layers-humans-can-read":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8072,"ai-vision-models-miss-hidden-text-layers-humans-can-read","AI Vision Models Miss Hidden Text Layers Humans Can Read","A new benchmark shows leading vision-language models reliably read only one of two overlapping text layers in an image, while humans read both.","AI vision models can miss half the words hiding in a single image, and a new benchmark just proved it with two fonts stacked on top of each other.\n\nResearchers built a 300-image dataset called DecoyBench using a method they call Decoy Font: each image pairs sharp, high-contrast text with a second layer of soft, blurred text sitting underneath it. Six closed-source vision-language models (VLMs) from three different model families were shown the images under two prompting setups - one that just asked what the image said, and one that explicitly told the model two layers were present - at both a normal resolution (512x512 pixels) and a tiny one (64x64 pixels). Human volunteers who looked at the same images read both layers of text with high accuracy, no matter the setup.\n\nThe models did not. At high resolution, most of them read the sharp-edged text about as well as humans do, but almost never pulled out the soft, blurred text underneath, even when told it was there. Shrink the image down and the pattern flips: the sharp text becomes unreadable to humans and machines alike, while the models suddenly read the blurred text just fine.\n\nThat flip is the real finding. It is not that these models are simply bad at reading blurry or overlapping text - they consistently favor one kind of visual detail over another, in a way humans do not. That is a structural blind spot baked into how the models process an image, not a resolution problem, and it shows up the same way across model families and prompting styles.\n\nAny product that leans on a VLM to read what is actually in an image - screenshots, scanned forms, content moderation queues - now has a documented, repeatable way to make it see only half the message.","[\"ai\",\"vision-language-models\",\"computer-vision\",\"research\"]","2026-09-28T04:00:00.000Z","2026-09-28T08:17:36.977Z","2026-09-28T08:17:43.320Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The CLIP typographic-attack anecdote is added without any citation or attribution and misstates the well-known real-world example — the famous sticker attack used the label \"iPod,\" not \"iPad\" — so either fix the product name and add a proper citation for this claim, or cut the anecdote entirely.","resolved","ai",[30,32,33,34],"vision-language-models","computer-vision","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.31403",0,{"sections":41},[42,45,49,54,59,64,68,73,78,83,88,93,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",4750,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",759,{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",261,"2026-09-27T15:30:35.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",188,"2026-09-27T20:46:36.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",151,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":94,"slug":95,"count":91,"latest_published_at":96},"General","general","2026-09-26T17:02:42.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]