[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-fix-for-a-basic-flaw-in-how-ai-agents-learn-from-mistakes":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},7701,"a-fix-for-a-basic-flaw-in-how-ai-agents-learn-from-mistakes","A Fix for a Basic Flaw in How AI Agents Learn From Mistakes","A new training framework called GRAFT traces credit through failed AI agent attempts instead of throwing away everything they got right.","A new reinforcement learning method aims to fix a basic flaw in how AI agents learn from multi-step tasks.\n\nGroup-based RL methods like GRPO have become the standard way to train reasoning and agentic large language models, and they work well when a task is a single response scored all at once. The problem shows up in multi-turn agent work, where GRPO judges an entire trajectory as good or bad rather than crediting individual steps. That means a failed run's one smart decision gets punished along with its mistakes, and a lucky win can reward steps that had nothing to do with the outcome. A new paper proposes GRAFT, a graph-based framework that stitches every sampled rollout into a single trajectory graph, recovers state values with Bellman iteration, and assigns credit to each step from the value difference between connected nodes. The authors also extend generalized advantage estimation to that graph structure, calling it Graph GAE, to cut down on noisy value estimates.\n\nThis matters because step-level credit assignment is the unglamorous plumbing behind every agent that calls tools, browses the web, or writes and runs code across multiple turns. Get it wrong and training either ignores good moves buried in bad runs or rewards good luck, both of which slow down or destabilize learning. As labs push agents into longer, tool-heavy workflows, getting per-step credit right matters more than another round of scaling.\n\nThe paper reports consistent gains over GRPO and other recent agentic RL methods on multi-turn benchmarks, though the code is only promised on GitHub, not yet published, so outside verification will have to wait.","[\"reinforcement learning\",\"ai agents\",\"llm training\",\"arxiv\"]","2026-09-25T04:00:00.000Z","2026-09-25T05:12:11.600Z","2026-09-25T05:12:17.810Z","published",null,[],"ai",[26,27,28,29],"reinforcement learning","ai agents","llm training","arxiv",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28963",0,{"sections":36},[37,40,45,50,55,60,65,70,75,80,85,90,94,99],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",4466,{"name":41,"slug":42,"count":43,"latest_published_at":44},"Security","security",729,"2026-09-24T19:54:21.000Z",{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",386,"2026-09-24T23:50:55.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",237,"2026-09-24T22:00:00.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Hardware","hardware",182,"2026-09-25T01:25:53.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Science","science",138,"2026-09-24T18:24:52.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Consumer Tech","consumer-tech",128,"2026-09-24T19:24:34.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Software","software",88,"2026-09-24T23:06:55.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",79,"2026-09-22T22:21:13.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Startups","startups",71,"2026-09-24T20:45:00.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Gaming","gaming",46,"2026-09-24T17:52:29.000Z",{"name":91,"slug":92,"count":88,"latest_published_at":93},"General","general","2026-09-25T02:12:57.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]