A new paper finds a free fix for a shortcut baked into the algorithms that train today's AI systems.
Most reinforcement-learning methods used to train AI models, including PPO (Proximal Policy Optimization), TRPO (Trust Region Policy Optimization), and GRPO (Group Relative Policy Optimization), estimate how well an updated version of the model would perform using data collected under the old version, rather than the new one. That substitution is what makes training affordable, but it introduces a bias that grows the further the new version drifts from the old one. That is why these methods clip updates or restrict them to a trust region, and why they cannot reuse a batch of old data once it is too far out of date. The researchers show that for certain setups, including standard token-by-token text generation, the missing correction term can be computed exactly from numbers the algorithms already calculate, at no extra cost.
That matters because sample efficiency is a real bottleneck in training large models. The exact fix is not free everywhere; it adds noise that grows with sequence length, so the authors built a dial that ranges from no correction (plain PPO) to full correction. Used briefly, early in training, when stale-data bias is worst, it sped up learning on difficult scheduling tasks, with bigger gains on harder problems. Left on the whole time, or used where clipping already covers the bias, it did nothing or made things worse.
This is not a replacement for PPO, just a sharper tool for a specific moment in training. Still, it is the kind of unglamorous, well-targeted insight that tends to quietly show up in training pipelines a year from now rather than in a press release.