AI/ ai-agents · benchmarks · reinforcement-learning · ai

CompoWorld Trains AI Agents to Chain Actions Across Services

CompoWorld teaches AI agents to chain actions across simulated software services, and its benchmark comparison cites an unverified Claude Opus 4.6.

A new paper proposes training AI agents by making them juggle dozens of software services at once, not just one app at a time.

Researchers behind CompoWorld built a library of 448 simulated software services exposing more than 10,000 tools, then used a random-walk process to chain those services together into multi-step tasks that require passing information from one service to another. Verified task completions were used to fine-tune a 35-billion-parameter model, Qwen3.6-35B-A3B, and a reinforcement-learning stage rewarded the agent for finishing every part of a task rather than just the easy parts. Across eight benchmarks, the trained agent beat its own untrained baseline by an average of 9.17 points. The paper also reports that on a benchmark called AutomationBench, its agent beats a model it names "Claude Opus 4.6" - a label that does not correspond to any Anthropic model we can verify, so that particular comparison should be read with caution.

Most agent-training environments still live inside a single sandboxed app, which trains a bot to be great at one tool and useless the moment a task spans two. Building a reusable library of services that can be randomly chained together is a cheaper way to manufacture the kind of cross-service, real-world busywork - check an order, then email a vendor, then update a spreadsheet - that people actually want automated.

A simulated dependency graph is still a simulation, though, and an agent that aces a benchmark built from fake services may have a rougher time the day a real API changes its schema - or the day someone checks what "Claude Opus 4.6" is actually supposed to mean.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →