AI/ ai agents · arc-agi-3 · self-improvement · benchmarks

AI Agent Clears Every ARC-AGI-3 Game by Writing Its Own Rules

Memento 3 lets a frozen LLM write and test its own world-model code, clearing all 25 ARC-AGI-3 games without updating its weights.

A frozen large language model just cleared every public game in the ARC-AGI-3 benchmark, without anyone updating its weights.

The system, called Memento 3, keeps a plain-language rulebook of hypotheses about how its environment works, then compiles that rulebook into executable code it uses to predict outcomes and plan moves. When a prediction fails, the agent rewrites the rulebook, recompiles, and accepts the new code only if it matches the rulebook's intent and reproduces the exact transitions it has already observed. The researchers report the single-model version solved all 25 public ARC-AGI-3 games, averaging 100.0 on a human-efficiency metric while using 44% of the actions a person would typically need. A separate test had the system learn a Pong controller that won 21-0 across three rounds with different openings, with no further calls to the underlying LLM.

The verification loop is the interesting part. Instead of fine-tuning the model itself, Memento 3 pushes the learning into an external, auditable record: a rulebook a human can read, paired with code that has to prove it matches observed reality before it is trusted. That is a meaningfully different shape of self-improvement than the usual weight-updating kind, and it directly targets a known problem: with only a few observations, multiple world models can fit the data equally well until more evidence arrives.

ARC-AGI-3's games are still small, discrete grids, and a 21-0 Pong score is a fun demo rather than proof this scales to messier, continuous, or adversarial settings, especially since the work is an unreviewed preprint that hasn't gone through peer review.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →