[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-compoworld-trains-ai-agents-to-chain-actions-across-services":10,"sections":49},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":39,"tags":40,"sources":44,"feedback":48,"feedback_at":22,"cost_usd":48,"total_tokens":48},9246,"compoworld-trains-ai-agents-to-chain-actions-across-services","CompoWorld Trains AI Agents to Chain Actions Across Services","CompoWorld teaches AI agents to chain actions across simulated software services, and its benchmark comparison cites an unverified Claude Opus 4.6.","A new paper proposes training AI agents by making them juggle dozens of software services at once, not just one app at a time.\n\nResearchers behind CompoWorld built a library of 448 simulated software services exposing more than 10,000 tools, then used a random-walk process to chain those services together into multi-step tasks that require passing information from one service to another. Verified task completions were used to fine-tune a 35-billion-parameter model, Qwen3.6-35B-A3B, and a reinforcement-learning stage rewarded the agent for finishing every part of a task rather than just the easy parts. Across eight benchmarks, the trained agent beat its own untrained baseline by an average of 9.17 points. The paper also reports that on a benchmark called AutomationBench, its agent beats a model it names \"Claude Opus 4.6\" - a label that does not correspond to any Anthropic model we can verify, so that particular comparison should be read with caution.\n\nMost agent-training environments still live inside a single sandboxed app, which trains a bot to be great at one tool and useless the moment a task spans two. Building a reusable library of services that can be randomly chained together is a cheaper way to manufacture the kind of cross-service, real-world busywork - check an order, then email a vendor, then update a spreadsheet - that people actually want automated.\n\nA simulated dependency graph is still a simulation, though, and an agent that aces a benchmark built from fake services may have a rougher time the day a real API changes its schema - or the day someone checks what \"Claude Opus 4.6\" is actually supposed to mean.","[\"ai-agents\",\"benchmarks\",\"reinforcement-learning\",\"ai\"]","2026-10-01T04:00:00.000Z","2026-10-02T04:30:58.764Z","2026-10-02T04:31:02.523Z","published",null,[24,30,35],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The closing skepticism claims AutomationBench 'is built around the researchers' own service library,' but the source never says this — AutomationBench appears to be one of the eight external benchmarks used for evaluation, not something derived from CompoWorld's 448-service library — so this unsupported claim must be corrected or removed.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The article references beating \"Claude Opus 4.6,\" but no such model exists in the Opus line (current Opus release is 4.8, while 4.6 corresponds to Sonnet), indicating an unverified or incorrect model name that needs correction before publication.",{"id":36,"reviewer":26,"round":37,"reason":38,"status":29},"editor-r3",3,"The draft silently swapped the source's claim from 'Claude Opus 4.6' to 'Claude Sonnet 4.6' without any attribution or flag — this is an invented\u002Funverified claim since the source paper never mentions Sonnet 4.6; instead, explicitly note that the paper claims to beat 'Claude Opus 4.6,' which does not match any real Opus release, and flag that discrepancy rather than substituting a different real model name.","ai",[41,42,43,39],"ai-agents","benchmarks","reinforcement-learning",[45],{"name":46,"url":47},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.33665",0,{"sections":50},[51,54,58,62,67,72,76,81,86,90,95,100,105,110],{"name":52,"slug":39,"count":53,"latest_published_at":18},"AI",5659,{"name":55,"slug":56,"count":57,"latest_published_at":18},"Security","security",818,{"name":59,"slug":60,"count":61,"latest_published_at":18},"Policy","policy",430,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":18},"Science","science",163,{"name":77,"slug":78,"count":79,"latest_published_at":80},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":87,"slug":88,"count":84,"latest_published_at":89},"Software","software","2026-09-30T21:41:11.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":111,"slug":112,"count":113,"latest_published_at":114},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]