Acing a chemistry exam and understanding chemistry may not be the same thing when the test-taker is a language model.
Researchers compared human student responses to a high school chemistry test and the quantitative reasoning section of a university entrance exam against responses from six multimodal LLMs answering the same items. Using exploratory factor analysis, factor congruence, and resampling, they checked whether the statistical relationships between questions and underlying skills looked the same for humans and machines. They found systematic differences in factor structure across both instruments. In plain terms, the patterns of which questions cluster together to reveal an underlying skill differ between the two groups, even when overall scores look comparable.
This matters because so much of the case for LLM capability rests on borrowed instruments: pass the bar exam, ace the SAT, and treat that score the way you would treat a human's. The study suggests that inference is shakier than it looks, since a model can hit the same score as a student while solving a structurally different problem underneath. That gap undermines any claim that a benchmark score tells you what a model actually understands, not just what it can output.
The next time a lab boasts that its model scored in the 99th percentile, it is worth asking what, exactly, that percentile was measuring.