Ask an AI to flip a coin, and it cheats in oddly human ways.
Researchers tested how large language models generate binary random sequences, using the classic simulated coin-flip paradigm from behavioral science. They compared single flips, 20-flip sequences, n-gram statistics, run lengths, alternation rates, and next-flip predictability against true random (Bernoulli) baselines and existing human data from prior studies. The models reproduced well-documented human biases: they alternate between heads and tails too often, avoid long streaks, and show a slight first-flip preference. Turning up the "temperature" setting, which adds more randomness to outputs, softened some rigid patterns but didn't erase the underlying structure, and follow-up tests, including prompt changes, continuation tests, corpus searches, and internal model probes, ruled out simple memorization or tokenization quirks as the cause.
That matters because plenty of real-world uses, like sampling data, simulating survey respondents, or generating test cases, assume an LLM can stand in for a coin flip or a human guess. It can't, reliably. That's a quiet but consequential caveat for anyone leaning on these models for anything resembling unbiased chance.
Humans have been bad at faking randomness since long before chatbots existed; the machines, it turns out, learned that particular flaw a little too well.