[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-build-ai-agents-that-edit-their-own-bad-reasoning":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},7576,"researchers-build-ai-agents-that-edit-their-own-bad-reasoning","Researchers Build AI Agents That Edit Their Own Bad Reasoning","A new world model lets AI agents rewrite their own flawed reasoning mid-task, tackling a failure mode researchers call task-state contamination.","AI agents that run long, multi-step tasks have a habit of doubling down on bad assumptions instead of correcting course. Researchers propose a fix that edits the agent's own reasoning instead of just predicting what happens next.\n\nMost \"world models\" built for AI agents work by forecasting environment responses, essentially guessing what a tool or webpage will return. The new Agent-Editing World Model (AEWM) skips that and instead tracks how an agent's reasoning and actions move a task forward. An \"Action Judge\" component sorts each decision into three buckets: critical, exploratory, or noisy. A \"State Revision\" component then rewrites the noisy ones using the history the agent already has, and a third piece called EditAct applies those rewrites mid-task, using real execution feedback rather than just flagging problems after the fact. The researchers trained and tested the system across three domains: web search, terminal use, and software engineering tasks.\n\nThe target problem, which the paper calls \"task-state contamination,\" is familiar to anyone who has watched an agent spiral: one bad guess early in a session quietly corrupts every decision that follows, and nothing in most agent architectures catches it. Editing the reasoning trail directly is a more targeted approach than the usual fix of just adding more memory or more retries.\n\nThe paper reports EditAct beating the strongest baseline by 3.2 to 6.7 points across six benchmarks and three agent backbones, plus a 70.5% macro-F1 score on an Action Judge benchmark the researchers built themselves. That last figure, scored on a test of their own design, is worth filing under \"claimed\" until someone outside the lab reproduces it.","[\"ai agents\",\"world models\",\"llm research\",\"ai benchmarks\"]","2026-09-24T04:00:00.000Z","2026-09-24T07:32:44.136Z","2026-09-24T07:32:50.382Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"publisher-r1","publisher",1,"The body reports 70.5% macro-F1 as beating the strongest baseline by 10.6 points, implying the baseline scored ~59.9%, but this precise comparison figure cannot be verified and reads as a potentially fabricated\u002Fover-specific statistic typical of unverifiable AI-generated research claims.","resolved","ai",[32,33,34,35],"ai agents","world models","llm research","ai benchmarks",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28416",0,{"sections":42},[43,46,50,55,60,65,70,75,80,85,90,95,100,105],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",4424,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",724,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",380,"2026-09-23T22:53:43.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",227,"2026-09-24T11:08:33.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",174,"2026-09-24T10:10:29.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",136,"2026-09-24T09:00:00.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",116,"2026-09-24T00:51:49.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Software","software",85,"2026-09-23T20:00:00.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Dev Tools","dev-tools",79,"2026-09-22T22:21:13.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",66,"2026-09-23T17:28:38.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",45,"2026-09-22T15:35:06.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",43,"2026-09-21T23:48:56.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",27,"2026-09-22T13:00:00.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]