AI/ ai · benchmarks · computer-vision · multimodal-models

Why AI Vision Benchmarks May Be Measuring Text, Not Sight

A new benchmark finds top AI vision models lose over half their accuracy once polar layouts remove the text-coordinate shortcut they rely on.

Leading multimodal AI models ace grid-based visual reasoning tests by quietly converting pictures into coordinate lists instead of actually looking at them.

Researchers built Polaris-Bench, a set of 53 visual reasoning tasks reformulated in polar coordinates, each paired with a Cartesian version that tests identical logic. Polar layouts break the neat rows-and-columns structure models exploit to translate an image into text coordinates. Across 14 state-of-the-art multimodal models, accuracy on standard Cartesian grids ranged from 69% to 83%. On the Polar versions of the same tasks, scores collapsed to 31-39%, and prompting tricks meant to boost reasoning barely moved the needle.

The gap suggests a chunk of what benchmarks call visual reasoning has really been text-based deduction wearing a picture's clothing. Human testers held 88.8% accuracy on the same polar tasks, which rules out the simple explanation that this is just a hard visual problem and points squarely at a shortcut baked into how these models process grids.

It's a useful gut check for anyone citing benchmark leaderboards as proof of visual intelligence: a lot of that score may belong to a model's grasp of coordinates, not its eyes.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →