A new benchmark finds that even the best AI models can't read a room the way a toddler can.
Researchers built Humanity's Sixth Sense (HSS), a benchmark of images and videos paired with human-written prompts that probe the kind of instant, implicit reasoning people do without thinking: guessing what happened just before a photo was taken, judging whether a car will fit in a parking spot, or picking up on who is in charge in a room. Human testers scored 93.1% accuracy on these tasks. The strongest model tried, GPT-6-astra, managed only 53.6%, even when run at maximum reasoning effort. Letting models actively manipulate the images and video in an agentic setup narrowed that gap, but did not close it.
The gap matters because this is exactly the kind of perception people lean on for everyday navigation and social interaction, not obscure trivia. Most AI benchmarks test either deliberate expert-level analysis, like math competition problems, or raw low-level perception, like spotting objects in a frame. HSS targets the fast, in-between skill of inferring cause, intent, and hidden structure from a glance, which is precisely what an assistant or robot working alongside people needs, and precisely where current models still fall short.
It is a useful check on the assumption that bigger models automatically get more human. A system that aces graduate-level exams but cannot tell whether two people in a photo are arguing is still missing something basic about how people actually see the world.