A new benchmark called Population Fidelity checks whether AI chatbots actually capture the range of human opinion, or just average it away.
Researchers built the framework around three measures: whether a model gets each group's answers roughly right, how much variation shows up between groups, and whether that variation lines up with the correct groups. They reran a prior study of "machine bias" in LLM survey responses, testing both the original models and newer ones. The pattern held up: poor representation came not just from too little variation between groups, but from variation assigned to the wrong groups entirely. They also tested cultural fine-tuning, a technique meant to pull model answers closer to local attitudes, and found it nudged models toward the overall survey average without improving how well they captured differences within the population.
That distinction matters because AI tools are increasingly pitched as stand-ins for survey panels and focus groups, a use case that only works if a model reflects real subgroup diversity rather than a flattened composite. Standard accuracy metrics, which look only at aggregate agreement, miss this failure mode entirely. A model can look well-calibrated on average while still getting the specifics of who holds which views wrong.
A chatbot that nails the national average while scrambling who actually thinks what isn't a cheap substitute for polling. It's a confident guess dressed up as data.