Leading multimodal AI models ace grid-based visual reasoning tests by quietly converting pictures into coordinate lists instead of actually looking at them.
Researchers built Polaris-Bench, a set of 53 visual reasoning tasks reformulated in polar coordinates, each paired with a Cartesian version that tests identical logic. Polar layouts break the neat rows-and-columns structure models exploit to translate an image into text coordinates. Across 14 state-of-the-art multimodal models, accuracy on standard Cartesian grids ranged from 69% to 83%. On the Polar versions of the same tasks, scores collapsed to 31-39%, and prompting tricks meant to boost reasoning barely moved the needle.
The gap suggests a chunk of what benchmarks call visual reasoning has really been text-based deduction wearing a picture's clothing. Human testers held 88.8% accuracy on the same polar tasks, which rules out the simple explanation that this is just a hard visual problem and points squarely at a shortcut baked into how these models process grids.
It's a useful gut check for anyone citing benchmark leaderboards as proof of visual intelligence: a lot of that score may belong to a model's grasp of coordinates, not its eyes.