AI/ ai · reinforcement-learning · robotics · game-ai

New Method Has an LLM Write the Controller, Not Play the Game

A new technique has an AI write reusable controller code instead of making live decisions, reacting faster than typical reinforcement learning policies.

A new AI system writes its own game-playing code, then throws away the language model before it ever hits play.

Researchers behind a project called Code to Control used a large language model to design the structure of a Python controller, then used a separate, derivative-free search method to tune its numeric parameters with feedback from the environment. Once tuning finishes, the controller runs on its own: no LLM calls, no step-by-step planning, just code executing directly as the policy. The team tested it across a batch of Atari games, Flappy Bird, and several MuJoCo robotic locomotion tasks. They report the resulting controllers react faster than Proximal Policy Optimization, or PPO, a widely used reinforcement learning algorithm that trains a neural network policy through repeated trial and error, and stay competitive with deep reinforcement learning baselines while using fewer environment interactions to get there.

Most LLM-based control setups either call the model at every single decision or lean on a learned world model that has to replan each step, and both approaches add latency that becomes a real problem outside of turn-based games. Code to Control sidesteps that by treating the language model as a one-time architect rather than a running brain, which may explain why it also kept working after researchers changed the environment's dynamics substantially, without a full retrain.

It is a useful reminder that not every control breakthrough needs a model running nonstop - sometimes a smarter one-time plan beats a faster loop.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →