A new paper argues that the real bottleneck in training AI agents through reinforcement learning isn't the algorithms, but the plumbing underneath them.
The researchers behind AgentFly built a framework that treats every sandbox, model service, and external API an agent touches as a typed, schedulable resource rather than a one-off script. The system splits into four layers: an agent layer for defining agents, tools, and reward functions; a rollout layer that runs the agent loops and scores outcomes; a context layer that organizes those rollouts and feeds in contextual information; and a resource layer that handles acquisition, reuse, and cleanup of the underlying infrastructure. The team trained agents across multiple tasks and models using this setup, and ran what they call the first controlled throughput comparison against other agentic reinforcement learning frameworks.
That throughput comparison is the detail worth watching. Agentic reinforcement learning has been creeping up as the default way to build capable AI agents, replacing prompt engineering and supervised finetuning, but the environments agents interact with are heterogeneous and expensive to spin up and tear down repeatedly. If resource scheduling really is the limiting factor on training scale, as the paper claims, then frameworks like this matter more than another round of algorithm tweaks.
Worth noting: the benchmark comparing throughput against rival frameworks comes from AgentFly's own authors, not a neutral third party, so treat the numbers as a starting point for scrutiny rather than a settled verdict.