[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-eyevqa-benchmark-exposes-gaps-in-eye-scan-ai-models":10,"sections":39},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":34,"feedback":38,"feedback_at":22,"cost_usd":38,"total_tokens":38},8843,"eyevqa-benchmark-exposes-gaps-in-eye-scan-ai-models","EyeVQA Benchmark Exposes Gaps in Eye Scan AI Models","A new ophthalmology benchmark shows the best AI vision model manages only an overall score of just 62.8, exposing gaps in spatial reasoning.","A new benchmark finds that even the best medical AI still struggles to read an eye scan.\n\nResearchers built EyeVQA from 21 public ophthalmic imaging datasets, combining them into 20,000 question-answer pairs across six disease groups and seven question formats, from simple true-false and multiple-choice to harder tasks like ranking severity, locating a point, and drawing a bounding box around a lesion. Answers come straight from the original clinical diagnoses, grading, and annotations rather than from another model, so scoring stays reproducible. Nearly half the questions require comparing multiple images at once, closer to how an ophthalmologist actually works. The team then tested 14 general-purpose, scientific, and medically specialized vision-language models in a zero-shot setup, with no fine-tuning allowed.\n\nThe top-performing model managed only an overall score of just 62.8, and every model did worse on spatial tasks like bounding boxes than on basic recognition questions. That gap matters because spatial grounding, the ability to point to exactly where a problem sits on a scan rather than just flagging that something looks off, is what a clinician actually needs from an assistive tool.\n\nVendors already market general AI models as able to read medical scans; this benchmark suggests that claim holds up far better for spotting a problem than for pinpointing it, which is the harder, more clinically useful trick.","[\"ai\",\"medical-ai\",\"benchmarks\"]","2026-09-30T04:00:00.000Z","2026-10-01T06:52:47.693Z","2026-10-01T06:52:54.172Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek states '62.8 percent accuracy' but the source only reports an 'overall score of 62.8' without defining it as an accuracy percentage — rephrase the dek to match the body's accurate wording ('overall score of just 62.8') instead of asserting an unsupported metric type.","resolved","ai",[30,32,33],"medical-ai","benchmarks",[35],{"name":36,"url":37},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.32352",0,{"sections":40},[41,45,50,55,60,65,70,75,80,84,89,94,99,104],{"name":42,"slug":30,"count":43,"latest_published_at":44},"AI",5270,"2026-10-01T04:00:00.000Z",{"name":46,"slug":47,"count":48,"latest_published_at":49},"Security","security",801,"2026-09-30T22:18:23.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",157,"2026-09-30T15:00:56.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":81,"slug":82,"count":78,"latest_published_at":83},"Software","software","2026-09-30T21:41:11.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]