AI/ ai · llms · surveys · research methods

A New Way to Catch AI Surveys Erasing Cultural Differences

A new diagnostic called CDP catches when AI-simulated survey panels flatten or exaggerate real cross-country differences that standard accuracy scores miss.

A new benchmark checks whether AI chatbots pretending to be survey respondents from different countries actually sound different from each other.

Researchers introduce Cultural Divergence Preservation (CDP), a diagnostic calibrated once against human data, to check whether cross-country differences in LLM-simulated survey answers are preserved, flattened, or exaggerated. They tested it across four LLM backbones, three persona-prompting methods, and two survey types - the World Values Survey and the Big Five personality test. Standard fidelity metrics like Jensen-Shannon divergence measure how close each country's simulated answers land to real data, but they miss whether the gaps between countries are being erased or overstated. CDP tracks that gap directly, and in controlled tests it moved predictably as researchers dialed cross-country differences up or down, while the standard metric barely budged.

The audit found that DeepPersona-Inspired prompting, a technique that scores well on conventional fidelity metrics, produced the worst flattening in every model and survey combination tested - meaning it made simulated respondents from different countries sound more alike than real survey takers actually are. That matters for anyone treating LLMs as a cheap substitute for expensive cross-national polling, since a method can look accurate on paper while quietly sanding down the cultural differences it's supposed to capture.

It's a useful reminder that a synthetic panel can nail the aggregate numbers and still miss the point: respondents who sound the same everywhere aren't simulating culture, they're erasing it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →