AI/ ai-agents · benchmarks · arc-agi-3 · llm-research

New AI Agent Learns Game Rules by Writing Its Own Code

Schema has an LLM agent turn messy trial-and-error into testable programs, pushing a key ARC-AGI-3 benchmark score from 58.7 percent to 99.2 percent.

A new agent harness called Schema teaches AI agents to learn the rules of unfamiliar environments by writing code, not prose.

Schema gives an off-the-shelf LLM agent a persistent workspace where it drafts programs, rather than prose notes, to describe how an unfamiliar environment behaves. The harness checks those programs against everything the agent has observed, lets it plan next moves inside the code, and verifies each step before the agent acts. Using the same base model, Schema pushed the ARC-AGI-3 RHAE score - Relative Human-Average Efficiency, a metric that benchmarks an agent's performance against typical human attempts - from 58.7 percent to 99.2 percent. It also cleared every public game in DiG-bench and matched the median score of the top 50 human players on MazeBench, a separate maze-solving test.

Most agent frameworks store what they have learned as free text, which is cheap to generate but hard to verify or reuse - the agent can contradict itself three screens later and never notice. Turning observations into executable programs forces consistency: a rule either holds against the interaction history or it does not. That is a meaningful jump for agents meant to operate in genuinely novel settings, like new software environments or games with no manual, rather than ones where the rules were baked into training data.

Whether that discipline survives outside benchmark games, where the rules are at least internally consistent, is the open question before anyone calls this a general-purpose skill.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →