A new benchmark wants to know if AI can actually solve puzzles, not just recognize patterns in them.
Researchers built PuzzleJAX, a GPU-accelerated engine that runs PuzzleScript games, the popular DSL that casual and professional designers have used to build puzzle games since 2013. Rather than hard-coding a fixed set of games like most GPU learning environments do, PuzzleJAX compiles any game written in the PuzzleScript language on the fly. The team validated several hundred of these community-made games, drawn from the thousands that exist in PuzzleScript's catalog, then ran tree search, reinforcement learning, and LLM reasoning models against them.
The interesting part is what the results show about the gap between "easy to understand" and "easy to solve." PuzzleScript games are simple enough that a human can grasp the rules in seconds, but the paper finds they often demand real planning and high-level insight, the kind of multi-step reasoning that trips up models that are otherwise good at pattern matching.
Most AI benchmarks either lean on a handful of iconic games, think Atari or Montezuma's Revenge, or use synthetic tasks built for tidy convergence curves. PuzzleJAX's hundreds of human-designed puzzles sit in a messier, more realistic middle ground, and that mess is probably the point.