A quiet bug in asynchronous AI agent training has been undermining the very math meant to keep that training stable.
Asynchronous reinforcement learning speeds up training for large language model agents by separating sample generation from policy updates. That split relies on breaking an importance-weighting term into two distinct pieces: one correcting for mismatches between inference-side and training-side probability distributions, and another limiting how far an update can stray from an older policy. The paper finds that real-world async pipelines, with their delayed updates and partial rollouts, routinely lose the historical logits needed to compute that split. Without those old logits, the two corrections tangle together, and the clipping and masking thresholds meant to keep training stable start interacting unpredictably.
That's a problem because asynchronous setups are becoming the default way labs scale reinforcement learning for LLM agents, trading precision for throughput. A correction mechanism that breaks quietly under realistic conditions, instead of failing loudly, can erode training runs for a long time before anyone traces the cause back to a missing logit.
The paper tests three exact fixes (snapshot-based version tracking, a dedicated old-logit model, and synchronization via partial rollout interruption) and one approximate route for when exact logits are too costly to keep. Its preferred method, a revised PPO-EWMA, reportedly delivers gains in both training speed and optimization performance.