Researchers have worked out the minimum number of simulated people an AI model actually needs to fairly represent a crowd's opinions.
The paper formalizes "simulation-augmented generation" (SAGE), a technique where a model generates stand-in personas representing different viewpoints before answering a contentious question, whether that's a political issue or a personal-advice dilemma. The authors borrow a fairness axiom from proportional-clustering research called mPJR+, which they describe as the strongest proportionality guarantee that centroid-based clustering can reliably satisfy. Using it, they prove a model doesn't need one simulated persona per real person in the population it's trying to represent. A much smaller set of simulated individuals is enough, and at answer time the model only needs to route a given prompt to a handful of those personas rather than consulting all of them. In tests on political questions and personal-advice queries, their routing method beat both k-means clustering and random selection at satisfying the mPJR+ fairness measure.
That efficiency gain is the real news here. Earlier pitches for SAGE-style systems implied that representing a population meant simulating a lot of it, which is slow and costly at inference time. This paper gives developers a mathematical basis for cutting that cost instead of just asserting that a model's answer reflects "diverse perspectives," which is usually where these claims stop.
Worth remembering: a proportionality guarantee only protects against a bad sampling algorithm, not a bad starting population. If the handful of simulated personas were built from a skewed or incomplete slice of real viewpoints to begin with, the math will faithfully and efficiently reproduce that skew.