AI/ llms · ai-research · synthetic-data · surveys

LLMs Agree With Themselves, Not With Each Other

A new study finds nine open-weight models give stable preferences alone but rarely agree with each other, undercutting their use as survey stand-ins.

A new study finds that large language models rarely agree on what people prefer, even though each one is remarkably consistent with itself.

Researchers tested nine open-weight models across three everyday choice domains - air travel, restaurants, and consumer products - asking each one to generate probability distributions over likely preferences. Repeated sampling showed each model's top answers stabilized quickly, so a given model is internally coherent. But stack the models against each other and the picture falls apart: there was little overlap even among the most probable outcomes, and the disagreement held up across temperature changes, greedy decoding, and tweaks to prompt wording and ordering.

That matters because researchers and companies increasingly use LLMs as cheap stand-ins for human survey respondents, on the assumption that a capable enough model approximates real population preferences. This study undercuts that assumption. It finds that which model you pick shapes the answer more than how you phrase the question - swap the model and the preferences shift more than if you'd just reworded the prompt.

So the synthetic focus group isn't measuring customer sentiment. It's measuring which lab trained the model doing the guessing.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →