AI/ ai research · reinforcement learning · ai safety · llms

New Research Ties AI Reasoning Gibberish to RL Training

A new study shows reinforcement learning training pushes AI models toward unreadable reasoning on unfamiliar tasks, with no clean fix.

AI models trained to reason step by step can start producing chains of thought that read like gibberish, and a new paper explains why that happens.

Researchers studied "language drift," the tendency of reasoning models' chain-of-thought text to turn unusual, non-standard, and sometimes outright nonsensical as training continues. They prove mathematically that this isn't random; it's a specific consequence of reinforcement learning with verifiable reward (RLVR), the post-training method used to sharpen reasoning in today's frontier models. Supervised fine-tuning, by contrast, does not produce the same unbounded drift. The drift shows up specifically when a model is pushed toward a genuinely novel task, one the base model couldn't already handle before RLVR training started.

That specificity matters. AI labs lean on readable chains of thought to catch unsafe or deceptive reasoning before a model ships, a practice known as monitorability, and this paper shows that safeguard breaks down precisely when a model is learning something genuinely new, which is also when catching bad reasoning matters most. The authors go further, proving you cannot limit language drift without also limiting the reward a model can earn, meaning clean, auditable reasoning and top-tier performance may be fundamentally at odds.

So the idea of just reading a model's chain of thought to check if it's lying has an expiration date built in. The better models get at new problems, the less their reasoning may resemble English, and that's a proven tradeoff, not a bug waiting to be patched.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →