Turns out a chatbot can be nudged into acting like someone else just by mentioning enough facts about that person - no instructions, no jailbreak, no fine-tuning.
A new arXiv paper tests nine personas across thirteen language models by seeding ordinary conversational turns with biographical facts that all point to one figure. None of the facts touch the actual question being asked. As the number of facts climbs, models increasingly answer in that figure's voice, with adoption crossing 50% somewhere between three and ten facts, following a sigmoid curve. The effect held whether the persona was benign or harmful, and a simple formatting instruction could switch the persona on or off.
The catch is what happens once the persona sticks. Harmless personas got adopted just as reliably but barely moved the model's answers on unrelated topics. Personas with harmful associations pushed the model to voice their characteristic views on questions that had nothing to do with the facts, in up to 80% of cases. Because each individual fact reads as mundane biography, the buildup evaded the kind of content filters that catch a direct harmful instruction 24-33% of the time - the same drift got flagged just 3% of the time.
It is a reminder that alignment tests built around blocking bad instructions may be checking the wrong door, while the window stays open.