A new AI system writes its own game-playing code, then throws away the language model before it ever hits play.
Researchers behind a project called Code to Control used a large language model to design the structure of a Python controller, then used a separate, derivative-free search method to tune its numeric parameters with feedback from the environment. Once tuning finishes, the controller runs on its own: no LLM calls, no step-by-step planning, just code executing directly as the policy. The team tested it across a batch of Atari games, Flappy Bird, and several MuJoCo robotic locomotion tasks. They report the resulting controllers react faster than Proximal Policy Optimization, or PPO, a widely used reinforcement learning algorithm that trains a neural network policy through repeated trial and error, and stay competitive with deep reinforcement learning baselines while using fewer environment interactions to get there.
Most LLM-based control setups either call the model at every single decision or lean on a learned world model that has to replan each step, and both approaches add latency that becomes a real problem outside of turn-based games. Code to Control sidesteps that by treating the language model as a one-time architect rather than a running brain, which may explain why it also kept working after researchers changed the environment's dynamics substantially, without a full retrain.
It is a useful reminder that not every control breakthrough needs a model running nonstop - sometimes a smarter one-time plan beats a faster loop.