A new paper proposes training AI agents by making them juggle dozens of software services at once, not just one app at a time.
Researchers behind CompoWorld built a library of 448 simulated software services exposing more than 10,000 tools, then used a random-walk process to chain those services together into multi-step tasks that require passing information from one service to another. Verified task completions were used to fine-tune a 35-billion-parameter model, Qwen3.6-35B-A3B, and a reinforcement-learning stage rewarded the agent for finishing every part of a task rather than just the easy parts. Across eight benchmarks, the trained agent beat its own untrained baseline by an average of 9.17 points. The paper also reports that on a benchmark called AutomationBench, its agent beats a model it names "Claude Opus 4.6" - a label that does not correspond to any Anthropic model we can verify, so that particular comparison should be read with caution.
Most agent-training environments still live inside a single sandboxed app, which trains a bot to be great at one tool and useless the moment a task spans two. Building a reusable library of services that can be randomly chained together is a cheaper way to manufacture the kind of cross-service, real-world busywork - check an order, then email a vendor, then update a spreadsheet - that people actually want automated.
A simulated dependency graph is still a simulation, though, and an agent that aces a benchmark built from fake services may have a rougher time the day a real API changes its schema - or the day someone checks what "Claude Opus 4.6" is actually supposed to mean.