[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-mllms-still-fail-at-finding-tiny-details-in-images":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},9309,"mllms-still-fail-at-finding-tiny-details-in-images","MLLMs Still Fail at Finding Tiny Details in Images","A new benchmark shows top multimodal AI models trail human accuracy by seven points when forced to actively search images for hidden evidence.","Researchers have built a benchmark that catches multimodal AI models faking their way through image analysis.\n\nVisualNeedle tests whether multimodal large language models actually look at images or just guess from context. The researchers found that prior benchmarks let models cheat three ways: inferring answers from wording alone, relying on coarse visual impressions instead of fine detail, and barely noticing when tools return corrupted image crops. VisualNeedle fixes this by hiding critical evidence in tiny, spatially isolated regions of information-dense scenes, and by adding a \"crop-black\" test that swaps real image crops for blank black ones to check if models actually use them. Across nine major MLLMs, text-only accuracy stayed below 10 percent and tool-free accuracy stayed below 20 percent. Even the best tool-using model hit just 56.00 percent, against a 63.00 percent human majority-vote baseline.\n\nThis matters because the industry has been selling \"over 90 percent accuracy\" scores on perception benchmarks as proof multimodal models see the way people do. VisualNeedle suggests those numbers largely measure pattern-matching and linguistic guesswork, not genuine visual search. That gap is exactly where real-world tasks like document review, surveillance footage analysis, or defect inspection would break current models.\n\nA benchmark that drops state-of-the-art accuracy from 90-plus to 56 percent isn't a rounding error -- it's a different test measuring a different skill, and the models aren't ready for it yet.","[\"multimodal ai\",\"benchmarks\",\"computer vision\",\"ai research\"]","2026-10-01T04:00:00.000Z","2026-10-02T08:39:38.608Z","2026-10-02T08:39:39.595Z","published",null,[],"ai",[26,27,28,29],"multimodal ai","benchmarks","computer vision","ai research",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.26380",0,{"sections":36},[37,40,44,48,53,58,62,67,72,76,81,86,91,96],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",5690,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",820,{"name":45,"slug":46,"count":47,"latest_published_at":18},"Policy","policy",430,{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",308,"2026-10-01T12:30:00.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":57},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",165,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",150,"2026-10-01T11:59:27.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":73,"slug":74,"count":70,"latest_published_at":75},"Software","software","2026-09-30T21:41:11.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]