A new study finds that large language models rarely agree on what people prefer, even though each one is remarkably consistent with itself.
Researchers tested nine open-weight models across three everyday choice domains - air travel, restaurants, and consumer products - asking each one to generate probability distributions over likely preferences. Repeated sampling showed each model's top answers stabilized quickly, so a given model is internally coherent. But stack the models against each other and the picture falls apart: there was little overlap even among the most probable outcomes, and the disagreement held up across temperature changes, greedy decoding, and tweaks to prompt wording and ordering.
That matters because researchers and companies increasingly use LLMs as cheap stand-ins for human survey respondents, on the assumption that a capable enough model approximates real population preferences. This study undercuts that assumption. It finds that which model you pick shapes the answer more than how you phrase the question - swap the model and the preferences shift more than if you'd just reworded the prompt.
So the synthetic focus group isn't measuring customer sentiment. It's measuring which lab trained the model doing the guessing.