[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-find-and-patch-a-flaw-in-async-ai-training":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8712,"researchers-find-and-patch-a-flaw-in-async-ai-training","Researchers Find and Patch a Flaw in Async AI Training","New research pinpoints why asynchronous RL training for AI agents breaks its own correction math, and fixes it with a revised PPO-EWMA method.","A quiet bug in asynchronous AI agent training has been undermining the very math meant to keep that training stable.\n\nAsynchronous reinforcement learning speeds up training for large language model agents by separating sample generation from policy updates. That split relies on breaking an importance-weighting term into two distinct pieces: one correcting for mismatches between inference-side and training-side probability distributions, and another limiting how far an update can stray from an older policy. The paper finds that real-world async pipelines, with their delayed updates and partial rollouts, routinely lose the historical logits needed to compute that split. Without those old logits, the two corrections tangle together, and the clipping and masking thresholds meant to keep training stable start interacting unpredictably.\n\nThat's a problem because asynchronous setups are becoming the default way labs scale reinforcement learning for LLM agents, trading precision for throughput. A correction mechanism that breaks quietly under realistic conditions, instead of failing loudly, can erode training runs for a long time before anyone traces the cause back to a missing logit.\n\nThe paper tests three exact fixes (snapshot-based version tracking, a dedicated old-logit model, and synchronization via partial rollout interruption) and one approximate route for when exact logits are too costly to keep. Its preferred method, a revised PPO-EWMA, reportedly delivers gains in both training speed and optimization performance.","[\"reinforcement-learning\",\"llm-training\",\"ai-research\",\"async-training\"]","2026-09-30T04:00:00.000Z","2026-09-30T21:04:45.304Z","2026-09-30T21:04:50.715Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Cut the closing line speculating that 'the authors have watched it happen more than once' — that's an unsupported inference not found in the source material; end instead on a concrete detail from the paper (e.g. the PPO-EWMA fix's stated benefit) rather than editorializing about the authors' experience.","resolved","ai",[32,33,34,35],"reinforcement-learning","llm-training","ai-research","async-training",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.12070",0,{"sections":42},[43,46,50,54,59,64,68,73,78,82,87,92,97,102],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5184,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",791,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",417,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",155,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]