AI/ vision-language models · ai benchmarks · compositional reasoning

New Benchmark Finds Vision AI Models Fail at Basic Binding

A new automated benchmark tests over 25 vision-language models and finds they still mix up which color or position belongs to which object.

Vision-language models still don't reliably know which color or position belongs to which object in a scene.

A team of researchers built an automated pipeline called Auto-Comp that generates matched pairs of test images for each concept: a bare, template-captioned version with objects floating on a white background, and a busier version with an LLM-rewritten caption placing the same objects in a realistic scene. That pairing lets the researchers isolate whether a model is actually binding attributes and relations to the right object, rather than guessing from keywords. They tested more than 25 models, spanning CLIP, SigLIP, models trained specifically to catch tricky mismatches, and frontier generative systems, across four tasks: color, shape-color, position, and relative size. Every model tested showed a large gap between catching an obvious swapped attribute and getting tripped up by similar-looking distractors, like two objects sharing a color.

That gap is the real finding. It shows the well-known "bag of words" problem, where models process captions as loose keyword soup, is only part of why these systems fail at basic reasoning; even models built to resist that specific failure still stumble on visual clutter. The paper also finds a trade-off worth watching: richer, more realistic scenes help models get spatial relationships right but make attribute binding worse.

So the fix for one flavor of confusion may sharpen another, which is a messier conclusion than the usual benchmark press release offers.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →