AI/ ai · computer-vision · eye-tracking · cognitive-science

AI Vision Models Find Targets Like Humans, Search Nothing Alike

Researchers found AI vision models locate search targets as well as humans, but their simulated eye movements skip the human scanning process entirely.

AI models can find what they're looking for almost as fast as we can. They just don't look for it the way we do.

A new study fed three general-purpose multimodal large language models the same foveated, human-matched visual input used in human eye-tracking experiments, then drove each model through the COCO-Search18 task fixation by fixation. On deciding whether a target was present and on reaching it, the models matched or beat human performance, detecting present targets near ceiling and landing on the target with their first saccade more often than people did. But the scanpaths themselves looked nothing like human ones: low-entropy, large-amplitude, and so self-consistent that a model agreed with its own repeated runs far more than two different humans agreed with each other. No amount of degrading the input closed that gap.

That's the tell. A model that reaches the right answer with an inhuman search pattern isn't modeling human vision, it's pattern-matching to the same endpoint through a single-pass shortcut. Standard alignment metrics, built around answers and saliency maps, would call this a success and miss the mismatch entirely.

Treat any AI vision benchmark that only scores outcomes with the same skepticism you'd apply to a resume that lists just the job title, not how the work got done.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →