AI/ computer-vision · vlm-judges · benchmarks · ai-evaluation

New Benchmark Shows AI Image Judges Can't Explain Their Picks

A new benchmark, VisionQ, tests whether vision-language models can justify image comparisons by criterion, not just pick a winner, and most fail.

AI judges are bad at explaining why one image beats another, a new benchmark finds.

VisionQ pulls 1,409 CVPR and ICCV papers and extracts more than 1,800 validated side-by-side comparison figures, the kind authors use to argue their method produces sharper, more consistent, or more realistic outputs than a rival's. Researchers hand-annotated 3,911 data points linking specific image crops to the exact visual claim the authors were making, then built a six-axis, 51-leaf taxonomy covering the different ways a vision model's output can be judged better or worse. The evaluation protocol strips out method names, captions, and paper identity, so a judge model has to pick the right output for the right stated reason, not just guess which one looks more polished. The team also trained VisionQ-Judge, a Gemma-4-E4B model tuned with DPO on matched evidence pairs, which cut "always pick the last option" bias by 7 percentage points and lifted accuracy by 2.5 points on held-out data.

Most benchmarks for AI-as-judge systems only ask "which image is better," which lets a model get credit for the right pick and the wrong reasoning. VisionQ closes that loophole, and the results aren't flattering: the best of 20 open- and closed-source VLM judges tested hit just 63.1% accuracy, against a 32.2% chance baseline, with accuracy swinging wildly depending on which visual criterion was being judged.

If "AI judged it" is becoming shorthand for objective evaluation in computer vision research, this is a reminder that the judging is still closer to an educated guess than a rubric.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →