AI/ llm-evaluation · cognitive-science · benchmarks · ai-research

New Benchmark Tests LLMs on Human Cognitive Skills, Not Tasks

A new benchmark built from neuropsychology tests shows LLMs and humans fail at different parts of the same tasks.

A new benchmark borrows tools from clinical psychology to find out what large language models actually can't do.

Researchers built NeuroCognition from three adapted neuropsychological tests: Raven's Progressive Matrices for abstract relational reasoning, a spatial working memory task for goal-directed spatial updating, and the Wisconsin Card Sorting Test for cognitive flexibility. Running it across 156 models, they confirmed a general factor of capability that shows up consistently on standard leaderboards. But performance drops sharply once tasks involve images instead of text, and drops further as complexity increases. Compared against a human baseline, models and people don't fail in the same spots - they stumble on different parts of the same puzzles.

That mismatch is the real finding. Most LLM benchmarks measure whether a model finishes a task, not whether it reasons the way a person does, so a model can top an image-reasoning leaderboard while still lacking the basic spatial updating a person handles without thinking. The researchers also found that throwing more complex reasoning at these tasks doesn't reliably help - simple, human-like strategies sometimes work better, which cuts against the assumption that more chain-of-thought is always the fix.

A model that aces trivia and still fumbles a test built for hospital patients is a good reminder that general capability and general intelligence are not the same claim.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →