A new AI system for real-time strategy games separates picking a game plan from executing it, and it wins more often because of that split.
Researchers built a reinforcement learning agent for MicroRTS, a stripped-down strategy game used as an AI testbed. The system splits into two layers: an executor trained with Proximal Policy Optimization that follows discrete commands covering economy, army composition, military posture and worker behavior, and a strategist that uses a Thompson-sampling bandit to pick which commands to issue based on what it observes the opponent doing. That strategist never identifies who it's playing against, it just watches and adapts. Tested against a single flat PPO policy trained with the same budget, architecture, curriculum and self-play league, the layered system won significantly more often against three of the four strongest opponents, including both held-out ones never seen in training, with win rates climbing as high as 0.97.
The point isn't the game itself. Deep RL agents are notoriously brittle outside their training distribution, a weakness that has dogged game-playing bots long after their headline-grabbing debuts. Decoupling what to do from how to do it is a plausible fix: it lets one well-trained execution policy get reused across strategies, rather than retraining a whole new brain every time an opponent breaks the pattern.
MicroRTS is a toy compared to StarCraft II or Dota 2, and a bandit swapping between a handful of preset command tuples is a long way from genuine strategic reasoning. Whether this approach holds up at real-game complexity is the actual question, and this paper doesn't answer it.