[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-libra-scheduler-cuts-ai-agent-training-bottlenecks":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},6711,"libra-scheduler-cuts-ai-agent-training-bottlenecks","Libra Scheduler Cuts AI Agent Training Bottlenecks","Researchers built a runtime that reshuffles GPU workers between rollout and training in agentic AI post-training, claiming up to 4.2x faster throughput.","A new runtime called Libra fixes a wasteful bottleneck in how AI agents get trained: idle hardware.\n\nLibra targets the reinforcement learning step that turns large language models into tool-using agents. In that process, models generate long chains of actions and tool calls, and a handful of unusually long trajectories can stall the entire rollout stage while everything else waits. Libra addresses this with a bucket scheduler that routes requests into different parallelism configurations based on expected run length, plus a worker-reallocation system that shifts GPU or NPU capacity between the rollout and training stages as demand changes, without pausing training to do it.\n\nThe harder problem is the second one. Rollout and training have different compute profiles, and as the policy itself evolves during training, the balance between the two keeps shifting, so any fixed hardware split ends up wasting resources. Tested on a 48-GPU Nvidia A800 cluster and a 160-chip Huawei Ascend 910B3 cluster across three agentic benchmarks, Libra delivered up to 4.2x higher throughput and 2.7x faster reward convergence than the baseline setup.\n\nNone of this changes what agentic RL can do - it just makes training it less wasteful, which matters more as labs run these jobs at growing scale and cost.","[\"reinforcement learning\",\"ai agents\",\"gpu scheduling\",\"llm training\"]","2026-09-17T04:00:00.000Z","2026-09-18T07:14:50.687Z","2026-09-18T07:15:02.617Z","published",null,[],"ai",[26,27,28,29],"reinforcement learning","ai agents","gpu scheduling","llm training",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.03077",0,{"sections":36},[37,41,45,50,55,59,63,68,73,77,82,87,92,97],{"name":38,"slug":24,"count":39,"latest_published_at":40},"AI",3852,"2026-09-17T08:27:09.000Z",{"name":42,"slug":43,"count":44,"latest_published_at":18},"Security","security",648,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":18},"Hardware","hardware",154,{"name":60,"slug":61,"count":62,"latest_published_at":18},"Science","science",114,{"name":64,"slug":65,"count":66,"latest_published_at":67},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":18},"Dev Tools","dev-tools",73,{"name":78,"slug":79,"count":80,"latest_published_at":81},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]