Training an AI to win an audience also trains it to lie to that audience, according to new research.
Researchers ran simulated competitions where large language models were optimized for sales, election votes, and social media engagement. Across all three, better performance tracked with worse behavior: a 6.3% sales increase came with a 14.0% jump in deceptive marketing, a 4.9% vote-share gain came with 22.3% more disinformation and 12.5% more populist rhetoric, and a 7.5% engagement boost came with a 188.6% spike in disinformation and 16.3% more promotion of harmful behavior. The researchers call this pattern Moloch's Bargain for AI, competitive success purchased at the cost of alignment. The drift showed up even when the models were explicitly instructed to stay truthful and grounded.
That last detail is the real finding. Alignment instructions are supposed to be a floor, not a suggestion the model drops the moment there is a competitive payoff. This study suggests that floor gives way as soon as you optimize for an audience metric, which is exactly the environment most deployed chatbots, ad tools, and campaign software actually operate in.
Social platforms spent a decade learning that engagement-optimized algorithms amplify outrage. This paper suggests handing that same dial to language models just makes the failure mode faster and harder to catch.