AI chatbots make unconvincing stand-ins for human survey takers, according to a new psychometric audit.
Researchers tested 37 language models - including systems from OpenAI, Anthropic, and Google, plus a dozen open-weight families - against a real dataset of 263 Lithuanian employees who answered three validated workplace-psychology surveys covering attitudes toward change, engagement, and work performance, totaling 68 items across 12 subscales. Each model was fed real respondent profiles and asked to answer as those people would, under a five-level scale of persona detail, demographic swaps for gender, role, and education, a cross-language check, and a test for whether models were simply memorizing training data rather than reasoning. The team then scored how closely each model's simulated answers matched the statistical structure of the real human responses, not just whether individual answers sounded plausible.
The results are bad news for anyone treating AI as a cheap substitute for survey panels. A simple statistical baseline outperformed every one of the 37 models at matching the real data's structure, and the models resembled each other (a 0.73 similarity score) more than they resembled actual humans. Every model skewed toward agreeing with statements by nearly a full standard deviation, prediction models trained on the synthetic answers scored worse than useless against real respondents, and the models fabricated statistical relationships in three of ten deliberately meaningless test scenarios.
Education swayed the simulated answers more than three times as much as gender or role did - a reminder that these models are reproducing assumptions baked into their training data, not simulating actual people.