[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-teach-ai-agents-to-rewrite-their-own-tools":10,"sections":46},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":41,"feedback":45,"feedback_at":22,"cost_usd":45,"total_tokens":45},8676,"researchers-teach-ai-agents-to-rewrite-their-own-tools","Researchers Teach AI Agents to Rewrite Their Own Tools","A new paper shows AI agents can evolve their own tooling faster by studying human demonstrations, not just their own trial-and-error game runs.","AI agents that modify their own code can now learn from watching a human play the game first.\n\nResearchers have described DemoEvolve, a system that lets a frozen large language model agent evolve the external program, or harness, that governs its behavior, using human demonstrations as a guide. Rather than relying only on its own costly trial-and-error rollouts, the agent examines demonstration recordings alongside its own rollout history, extracts the strategies used and the conditions under which they apply, then turns those into reusable pieces of its harness. The team tested the approach on two long-horizon games, Balatro and Slay the Spire 2, comparing demonstration-guided evolution against a self-rollout-only baseline and a version augmented with retrieved text-based knowledge, all under the same interaction budget. On held-out seeds, DemoEvolve raised mean capped final-round progress in Balatro from 16.83 to 20.00 and mean floor reached in Slay the Spire 2 from 18.17 to 28.83.\n\nLong-horizon agent tasks are expensive to run, so every rollout counts, and sparse feedback makes it hard for an agent to know which fix actually helped. Demonstrations give the agent a shortcut: concrete examples of what works, rather than forcing it to reverse-engineer strategy from delayed win-or-loss signals alone. That is a data-efficiency argument, not a raw-capability one, since the underlying model stays frozen throughout.\n\nCard games make convenient benchmarks because progress is easy to score, but it is a long way from shaving turns off a roguelike run to an agent that can usefully rewrite its own tooling on a messy real-world software task.","[\"ai-agents\",\"llm\",\"coding-agents\",\"benchmarks\"]","2026-09-30T04:00:00.000Z","2026-09-30T18:50:39.731Z","2026-09-30T18:50:46.035Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Add explicit attribution to the source (the arXiv preprint and its posting date) rather than the vague 'a new technique'\u002F'researchers,' and briefly explain what 'mean capped progress' and 'mean floor reached' actually measure so the 16.83→20.00 and 18.17→28.83 improvements are interpretable to readers.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The arXiv ID (2605.24539) indicates a May 2026 submission, contradicting the claim that it was 'posted to arXiv this week' relative to the September 30, 2026 date.","ai",[37,38,39,40],"ai-agents","llm","coding-agents","benchmarks",[42],{"name":43,"url":44},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.24539",0,{"sections":47},[48,51,55,59,64,69,73,78,83,87,92,97,102,107],{"name":49,"slug":35,"count":50,"latest_published_at":18},"AI",5184,{"name":52,"slug":53,"count":54,"latest_published_at":18},"Security","security",791,{"name":56,"slug":57,"count":58,"latest_published_at":18},"Policy","policy",417,{"name":60,"slug":61,"count":62,"latest_published_at":63},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":18},"Science","science",155,{"name":74,"slug":75,"count":76,"latest_published_at":77},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":88,"slug":89,"count":90,"latest_published_at":91},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]