Vision language models ace plenty of image benchmarks, but they still trip over simple abstract reasoning puzzles. Now a new paper explains why.
The researchers borrowed a test called Relational Match-to-Sample from psychology, where human toddlers and animals are judged on whether they match objects by shared features or by shared relationships. They ran it on frontier API models (GPT, Claude, Gemini) and three open-source families (Qwen3.5, Gemma-4, InternVL3), then dissected the models layer by layer. Four factors pushed models toward the correct relational answer: a stronger capability tier, bigger scale, fewer objects crammed into a scene, and less per-object visual noise. That progression looks a lot like the "relational shift" developmental psychologists see in human kids as they age.
Peering inside the models, the team found two circuits fighting for control: an early one that sorts images by surface-level object features, and a later one that sorts by abstract relationships. When they switched off the relational heads they had identified, performance on the unrelated ARC-AGI-1 benchmark dropped more than when they switched off random heads, meaning this isn't some quirk of their custom test. The same relational circuitry gets recruited elsewhere.
That's a real contribution, because most VLM failure reports stop at "the model got it wrong" without saying which cognitive step broke. This paper says it's not that the models can't see objects, it's that an early feature-matching circuit keeps winning the argument with a later relational one. For a field that throws more parameters and training data at every problem, that's a useful, humbler diagnosis.
It is still just a diagnosis, not a fix. Nobody has shown how to strengthen the relational circuit without also brute-forcing scale, which is the expensive lever every lab already pulls.