AI/ llm · fintech · ai-research · simulation

Study Finds LLMs Cannot Predict Individual Trading Behavior

An 80-person paper-trading study finds large language models fail to beat a simple recent-activity baseline at predicting what investors will trade next.

A new study tested whether large language models can actually simulate how real investors trade. They mostly cannot.

Researchers ran an 80-participant paper-trading experiment, tracking simulated transactions, portfolio states, and market data over time. LLMs were asked to predict, day by day, whether each user would trade next. Across 1,239 aligned user-days, no model reliably beat a bare-bones baseline that just assumes people keep doing what they did recently. Accuracy got worse from there: the models struggled even more to predict buy-versus-sell structure and which specific assets a user would pick, and small differences in predicted activity could snowball into very different portfolio outcomes.

This matters because LLM-based user simulators are being pitched as stand-ins for real people in product testing, trading-app design, and behavioral research, on the assumption that a good-enough language model can approximate an individual's decision-making. This study suggests that assumption does not hold yet. The models pick up short-term patterns, like the fact that someone researching a ticker heavily is likely to trade soon, but the researchers found no evidence that link is causal, and the models never recover anything resembling a stable, individual decision process.

It is a useful reality check in a year when "AI agent simulates a user" has become a pitch unto itself. Predicting that someone who traded yesterday will trade again tomorrow is not intelligence, it is momentum, and right now the fancier models are not clearing that low bar by much.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →