AI agents that modify their own code can now learn from watching a human play the game first.
Researchers have described DemoEvolve, a system that lets a frozen large language model agent evolve the external program, or harness, that governs its behavior, using human demonstrations as a guide. Rather than relying only on its own costly trial-and-error rollouts, the agent examines demonstration recordings alongside its own rollout history, extracts the strategies used and the conditions under which they apply, then turns those into reusable pieces of its harness. The team tested the approach on two long-horizon games, Balatro and Slay the Spire 2, comparing demonstration-guided evolution against a self-rollout-only baseline and a version augmented with retrieved text-based knowledge, all under the same interaction budget. On held-out seeds, DemoEvolve raised mean capped final-round progress in Balatro from 16.83 to 20.00 and mean floor reached in Slay the Spire 2 from 18.17 to 28.83.
Long-horizon agent tasks are expensive to run, so every rollout counts, and sparse feedback makes it hard for an agent to know which fix actually helped. Demonstrations give the agent a shortcut: concrete examples of what works, rather than forcing it to reverse-engineer strategy from delayed win-or-loss signals alone. That is a data-efficiency argument, not a raw-capability one, since the underlying model stays frozen throughout.
Card games make convenient benchmarks because progress is easy to score, but it is a long way from shaving turns off a roguelike run to an agent that can usefully rewrite its own tooling on a messy real-world software task.