AI/ vision-language models · cross-modal associations · eye tracking · arxiv

AI Models Match Human Choices on Bouba-Kiki, Not Human Attention

A new study finds vision-language models can match human word-shape choices but look at images differently than humans do, even after training on human data.

Vision-language models can guess the same 'bouba' or 'kiki' shape a human would pick, but they're not looking at the image the same way a human does.

A new arXiv preprint (2609.36475, posted September 30, 2026) tested this directly. Researchers showed 53 human participants and a set of vision-language models the same stimuli: a made-up word paired with two images, one round and one spiky, and recorded both choices and eye movements. The larger VLMs picked the same shape as human participants at above-chance rates, but where the models focused their visual attention lined up with actual human gaze worse than a simple fixed baseline that just assumes people look at the center of an image. Fine-tuning smaller VLMs on human choice data pushed their choice accuracy up to match a human majority vote on words and images the models hadn't seen before, yet their attention maps still trailed that same center-bias baseline. Training attention directly on recorded human gaze data improved the correlation between model attention and human gaze, but did nothing for choice accuracy.

This matters because AI evaluation increasingly leans on behavioral matching: if a model answers like a person, it's often treated as a proxy for human-like understanding. This paper shows that assumption breaks down under inspection, since two systems can converge on identical decisions through entirely different internal processes. That's a meaningful caution for anyone building or citing benchmarks that grade models purely on final answers rather than how those answers get made.

A model that reaches your answer isn't the same as a model that thinks the way you do, and this study is a data point worth remembering next time a benchmark claims human-level alignment.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →