[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-giving-ai-agents-code-skills-triples-game-progress":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8206,"giving-ai-agents-code-skills-triples-game-progress","Giving AI Agents Code Skills Triples Game Progress","A NetHack study finds letting AI agents call reusable code skills instead of raw actions nearly triples progress and cuts inference costs by 86 percent.","Researchers found that giving AI agents prewritten code skills - instead of forcing them to pick every low-level action - nearly triples how far they get in NetHack, a notoriously brutal roguelike.\n\nThe team built CodeHack, a library of NetHack skills paired with plain-language descriptions, then tested agents that could only use raw primitive actions against agents that could call skills, or mix skills with primitives as needed. In zero-shot testing across a broad set of NetHack scenarios, skill-equipped agents nearly tripled game progression compared to primitive-only agents, while cutting inference cost per episode by 86 percent. Agents that combined skills with the option to fall back on primitives kept most of that gain and stayed flexible for situations the skill library couldn't handle. In reinforcement learning, skill-based agents learned 7.2 times faster than primitive-only agents, measured by average gain in dungeon level over the same training budget.\n\nNetHack is a stand-in for the kind of long, multi-step task that trips up most language agents: no single prompt covers a full playthrough, and every extra low-level decision adds cost and error. Packaging that knowledge as callable code, rather than asking a model to re-derive it turn by turn, is a cheap way to buy both speed and accuracy - and the RL results suggest it also speeds up how fast an agent gets better with practice, not just how well it performs out of the box.\n\nIt's the same logic that got human programmers to stop writing everything in assembly: let a library handle the routine work, and save the reasoning for judgment calls the library can't make.","[\"ai-agents\",\"nethack\",\"reinforcement-learning\",\"code-generation\"]","2026-09-28T04:00:00.000Z","2026-09-28T16:52:23.999Z","2026-09-28T16:52:30.536Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The body lists supervised fine-tuning as one of three tested setups but never reports any SFT result, leaving a dangling reference — either cut SFT from that setup list or state its finding so all three named settings are followed through.","resolved","ai",[32,33,34,35],"ai-agents","nethack","reinforcement-learning","code-generation",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.31076",0,{"sections":42},[43,46,50,55,60,65,69,74,79,84,89,94,98,103],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",4844,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",762,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",265,"2026-09-28T14:00:00.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",189,"2026-09-28T10:52:40.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":18},"Science","science",151,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":95,"slug":96,"count":92,"latest_published_at":97},"General","general","2026-09-26T17:02:42.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]