AI/ ai · ab-testing · llm agents · statistics

Researchers Use LLM Personas to Cut A/B Testing Costs

A new statistical framework lets AI persona simulations stand in for some real test subjects, cutting A/B testing costs without sacrificing rigor.

A/B testing is slow and expensive. A new paper argues you can shrink it by letting AI personas predict some of the results before you run the real experiment.

The framework, laid out in a newly published research paper, uses machine-learning predictions to reduce the sample size a valid experiment needs, without pretending those predictions are flawless. It handles two kinds of forecasts: coarse, population-level signals that only guess the direction of an effect, and fine-grained, individual-level estimates. For the coarse case, the researchers use an asymmetric statistical test they prove stays consistent and robust even when the signal is weak. For fine-grained predictions, they introduce a method called Generalized PPI++ (GPPI), which extends an existing technique, Prediction-Powered Inference, to handle prediction errors that behave in nonlinear ways.

The predictions themselves come from LLM agents assigned user personas that simulate how a specific type of person would behave, tested across four real-world datasets. That's the actual pitch: instead of recruiting test subjects, you prompt a model to role-play as one. The paper reports the approach substantially cuts experimental costs while keeping the statistics valid, whether the AI predictions turn out to be accurate or not.

It's a hedge dressed up as a framework: the math is built to survive AI personas being wrong, which is itself a tacit bet that they often will be.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →