[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-method-has-an-llm-write-the-controller-not-play-the-game":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8891,"new-method-has-an-llm-write-the-controller-not-play-the-game","New Method Has an LLM Write the Controller, Not Play the Game","A new technique has an AI write reusable controller code instead of making live decisions, reacting faster than typical reinforcement learning policies.","A new AI system writes its own game-playing code, then throws away the language model before it ever hits play.\n\nResearchers behind a project called Code to Control used a large language model to design the structure of a Python controller, then used a separate, derivative-free search method to tune its numeric parameters with feedback from the environment. Once tuning finishes, the controller runs on its own: no LLM calls, no step-by-step planning, just code executing directly as the policy. The team tested it across a batch of Atari games, Flappy Bird, and several MuJoCo robotic locomotion tasks. They report the resulting controllers react faster than Proximal Policy Optimization, or PPO, a widely used reinforcement learning algorithm that trains a neural network policy through repeated trial and error, and stay competitive with deep reinforcement learning baselines while using fewer environment interactions to get there.\n\nMost LLM-based control setups either call the model at every single decision or lean on a learned world model that has to replan each step, and both approaches add latency that becomes a real problem outside of turn-based games. Code to Control sidesteps that by treating the language model as a one-time architect rather than a running brain, which may explain why it also kept working after researchers changed the environment's dynamics substantially, without a full retrain.\n\nIt is a useful reminder that not every control breakthrough needs a model running nonstop - sometimes a smarter one-time plan beats a faster loop.","[\"ai\",\"reinforcement-learning\",\"robotics\",\"game-ai\"]","2026-10-01T04:00:00.000Z","2026-10-01T09:30:53.793Z","2026-10-01T09:30:59.262Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Define the PPO acronym (Proximal Policy Optimization) on first use before citing 'faster than PPO' and 'competitive with deep reinforcement learning baselines' comparisons.","resolved","ai",[30,32,33,34],"reinforcement-learning","robotics","game-ai",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38733",0,{"sections":41},[42,45,50,55,60,65,70,75,80,84,89,94,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5350,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Security","security",801,"2026-09-30T22:18:23.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",157,"2026-09-30T15:00:56.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":81,"slug":82,"count":78,"latest_published_at":83},"Software","software","2026-09-30T21:41:11.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]