AI agents can now spawn copies of themselves to break down problems too big for one context window.
A new paper describes Recursive Agent Optimization (RAO), a reinforcement learning method for training what the authors call recursive agents, models that can spawn new instances of themselves and delegate sub-tasks to those instances, recursively. RAO teaches a model when to delegate, what to hand off, and how to combine results from its own clones. That delegation functions as an inference-time scaling trick: instead of stuffing more context into a single agent, the system divides a hard problem across a tree of smaller, focused agents. The authors report these trained agents handle tasks that exceed the base model's context window, generalize to problems much harder than their training tasks, and finish faster in wall-clock terms than single-agent systems.
Context-window limits are one of the main reasons engineers currently wire together multi-agent systems by hand, with routing logic nobody fully trusts. Training a model to learn its own delegation policy, instead of hard-coding it, could turn recursive multi-agent setups into a standard technique rather than a bespoke workaround, and the reported efficiency and speed gains over single-agent baselines suggest real practical upside, not just a cleverer architecture diagram.
One caveat: this is a v2 cross-listed arXiv submission, not a peer-reviewed benchmark result, so "generalizes to much harder tasks" is a claim worth revisiting once it's tested outside the authors' own evaluation.