[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-teach-ai-agents-to-plan-act-and-grade-themselves":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},6802,"researchers-teach-ai-agents-to-plan-act-and-grade-themselves","Researchers teach AI agents to plan, act and grade themselves","UnifiedPlayers lets an AI agent write its own tasks, solve them with code, and grade its own work, beating prior self-training methods on reasoning tests.","A new training framework has three AI agents write, solve, and grade their own practice problems - and the results beat systems that try to do all three jobs with one model.\n\nResearchers built UnifiedPlayers, a system with three specialized roles: a Planning Player that invents tasks, an Execution Player that solves them by writing and running Python code across multiple steps, and an Evaluation Player that builds automated checkers to grade the answers. All three train together using a reinforcement-learning method called GRPO, with rewards tuned so the roles improve in sync rather than working at cross-purposes. The team tested the setup on two different base models across twelve reasoning benchmarks. UnifiedPlayers beat the best prior method by at least 3.5% on math reasoning and 3.9% on general reasoning tasks.\n\nThe real innovation is the grading system. Prior self-training approaches leaned on fixed verifiers that go stale as models find new ways to fail, or on 'self-consistency' checks that can rubber-stamp a wrong answer just because a model reaches it several times. UnifiedPlayers' learned verifier catches adversarial wrong answers 84.2% of the time and produces a reward signal with more than twice the discriminating power of a self-consistency baseline, making it harder to fool.\n\nThat's a real fix for a real bottleneck in training AI agents without armies of human labelers - though it's worth remembering this is an arXiv preprint reporting benchmark scores, not a system anyone's agent is running in production yet.","[\"ai\",\"ai-agents\",\"reinforcement-learning\",\"research\"]","2026-09-18T04:00:00.000Z","2026-09-18T17:22:52.408Z","2026-09-18T17:23:04.323Z","published",null,[],"ai",[24,26,27,28],"ai-agents","reinforcement-learning","research",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.20089",0,{"sections":35},[36,39,43,48,53,57,61,66,70,75,80,85,90,95],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",4017,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",653,{"name":44,"slug":45,"count":46,"latest_published_at":47},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":18},"Hardware","hardware",155,{"name":58,"slug":59,"count":60,"latest_published_at":18},"Science","science",121,{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":18},"Dev Tools","dev-tools",77,{"name":71,"slug":72,"count":73,"latest_published_at":74},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]