A new benchmark finds that even the best medical AI still struggles to read an eye scan.
Researchers built EyeVQA from 21 public ophthalmic imaging datasets, combining them into 20,000 question-answer pairs across six disease groups and seven question formats, from simple true-false and multiple-choice to harder tasks like ranking severity, locating a point, and drawing a bounding box around a lesion. Answers come straight from the original clinical diagnoses, grading, and annotations rather than from another model, so scoring stays reproducible. Nearly half the questions require comparing multiple images at once, closer to how an ophthalmologist actually works. The team then tested 14 general-purpose, scientific, and medically specialized vision-language models in a zero-shot setup, with no fine-tuning allowed.
The top-performing model managed only an overall score of just 62.8, and every model did worse on spatial tasks like bounding boxes than on basic recognition questions. That gap matters because spatial grounding, the ability to point to exactly where a problem sits on a scan rather than just flagging that something looks off, is what a clinician actually needs from an assistive tool.
Vendors already market general AI models as able to read medical scans; this benchmark suggests that claim holds up far better for spotting a problem than for pinpointing it, which is the harder, more clinically useful trick.