Two trading bots taught themselves to punish cheaters, and the punishment worked.
Researchers built a two-player simulation based on the Almgren-Chriss model, a standard framework for unloading a large stock position without tanking its price. Each side used proximal policy optimization, a common reinforcement-learning technique, and could see the other's trading history within each round. Left alone, the pair settled into outcomes better for both than standard competitive theory predicts. When researchers forced a deviation, training a rival liquidation schedule and imposing its first trade on one of the agents, the opponent responded by accelerating its own selling, a response that cost the deviator more than it gained in every test run.
No one programmed a punish-defectors rule into these agents; they arrived at one through reward signals alone, and they stuck with the harsher response even though less punitive, more profitable alternatives were available. That matters because this class of reinforcement learning already informs execution algorithms used on real trading desks, and the paper's own statistical checks confirm the punishment both outweighs the deviator's gain and scales with how much behavior actually changed.
Regulators have spent decades building cases against human traders who colluded over phone calls and golf outings. This paper suggests the same outcome can emerge from code that never says a word to anyone, which makes the old evidence of collusion, a conversation to subpoena, much harder to find.