AI/ ai evaluation · llms · education · research methods

LLM Exam Scores May Not Mean What They Mean for Humans

A new study finds LLMs and humans get similar exam scores via different underlying skill structures, questioning what AI benchmarks really measure.

Acing a chemistry exam and understanding chemistry may not be the same thing when the test-taker is a language model.

Researchers compared human student responses to a high school chemistry test and the quantitative reasoning section of a university entrance exam against responses from six multimodal LLMs answering the same items. Using exploratory factor analysis, factor congruence, and resampling, they checked whether the statistical relationships between questions and underlying skills looked the same for humans and machines. They found systematic differences in factor structure across both instruments. In plain terms, the patterns of which questions cluster together to reveal an underlying skill differ between the two groups, even when overall scores look comparable.

This matters because so much of the case for LLM capability rests on borrowed instruments: pass the bar exam, ace the SAT, and treat that score the way you would treat a human's. The study suggests that inference is shakier than it looks, since a model can hit the same score as a student while solving a structurally different problem underneath. That gap undermines any claim that a benchmark score tells you what a model actually understands, not just what it can output.

The next time a lab boasts that its model scored in the 99th percentile, it is worth asking what, exactly, that percentile was measuring.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →