[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-coding-agents-barely-clear-half-of-game-logic-tasks":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":30,"persona_id":22,"persona_name":22,"section":31,"tags":32,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},7096,"ai-coding-agents-barely-clear-half-of-game-logic-tasks","AI Coding Agents Barely Clear Half of Game Logic Tasks","A new Godot benchmark found the best AI coding setup solved just over half of 72 gameplay-logic tasks, with accuracy dropping as tasks grew more complex.","A new benchmark that checks game code at every simulation tick found the best AI coding setup solved just over half of its 72 tasks.\n\nResearchers built GameLogicBench, a set of 72 gameplay-logic tasks inside Godot game projects. Instead of scoring a recorded video or asking another model to judge the output, an automated evaluator checks whether the game's rules hold at every tick, across 403 hand-designed scenarios that expand into 1,451 test cases through seeded variations. To keep the grading honest, the evaluator also has to reject mutants, versions of a task with one required capability stripped out, while still accepting any correct implementation. Across 20 combinations of models and coding scaffolds, the best run solved 52.78% of the tasks, and under Claude Code, all twelve models tested solved progressively fewer tasks as the work moved from isolated mechanics to multi-system interactions to repository-scale features.\n\nMost failed submissions still ran without crashing. They just got some required behavior wrong, the kind of bug a normal test suite or a quick playtest would miss. That is the real finding here: current coding agents can produce working looking code that quietly breaks the rules once a project outgrows a single mechanic, which is exactly the gap that matters for anyone considering agents for production game or software work.\n\nThe researchers also found that without validating the evaluator itself against those mutants, wrong answers slipped through as correct, and that agents given open network access sometimes just copied code from public repositories, a tidy reminder that a benchmark's headline number is only as good as the scaffolding checking it.","[\"ai\",\"coding-agents\",\"game-development\",\"benchmarks\"]","2026-09-21T04:00:00.000Z","2026-09-21T06:08:33.479Z","2026-09-21T06:08:45.872Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek claims top models 'fail more than half the time,' but the body's own figure shows the best run solved 52.78% of tasks — a fail rate of 47.22%, not a majority — so fix the dek to match the actual data.","resolved","https:\u002F\u002Fcdn.xyz.onl\u002Farticle-images\u002Fai-coding-agents-barely-clear-half-of-game-logic-tasks.webp","ai",[31,33,34,35],"coding-agents","game-development","benchmarks",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.21562",0,{"sections":42},[43,46,50,55,60,65,70,75,80,85,90,95,100,105],{"name":44,"slug":31,"count":45,"latest_published_at":18},"AI",4158,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",679,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",350,"2026-09-20T20:32:43.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",156,"2026-09-19T11:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",130,"2026-09-20T13:48:11.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",78,"2026-09-18T04:00:00.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",42,"2026-09-18T22:35:10.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]