A new trust-region trick could let AI labs train language models faster without the instability that usually comes with speed.
Researchers tackled asynchronous reinforcement learning, a setup where a model generates practice attempts while simultaneously updating itself, rather than waiting for each batch to finish before training on it. That overlap speeds things up, but the data goes stale and can destabilize training, sometimes causing a full policy collapse. Prior fixes clipped any token whose probability ratio strayed too far from the original policy, using the same cutoff everywhere. The authors argue that cutoff is wrong: the ratio's natural range depends on how uncertain the model was about that token, so a single threshold punishes legitimate exploration in uncertain spots while letting noise through in near-certain ones. Their fix, an entropy-normalized trust region called ENTR, adjusts the threshold position by position instead.
On long multi-step tasks and math benchmarks, ENTR beat existing asynchronous methods, including a 6.9% gain over the strongest baseline on BrowseComp-Plus, a benchmark for long, multi-step web research tasks. It kept training stable even when the data was up to 30 policy versions out of date, and matched a slower, fully synchronous baseline method while running 2.6 times faster. For anyone post-training large models, that's a real efficiency gain, not a rounding error.
The paper is a v5 update to an earlier preprint, not yet peer reviewed, so the headline numbers are promising rather than settled. Still, a 2.6x speedup with no apparent stability cost is exactly the kind of detail that gets copied into production training pipelines well before anyone finishes reviewing it.