A new benchmarking framework puts three very different flavors of AI puzzle-solving on the same scorecard, and the LLMs don't look great on cost.
Researchers built a three-dimensional framework that maps AI decision-making methods onto the Markov decision process (MDP) formalism, ranks how much human-designed structure each method leans on, and measures the skill and computational cost each one burns through. They tested it on the Tower of Hanoi, a puzzle with simple rules but complexity that scales fast as you add disks. Four approaches went head to head: Neurosolver, a graph-based solver; forward-backward reinforcement learning (FBRL); automated thought-of-search (AutoToS), an LLM-based method; and a two-agent variant called DA-ToS. The framework let the team compare all of them under the same conditions instead of relying on separate papers' cherry-picked benchmarks.
The standout finding is that LLM-based solvers don't actually encode less human knowledge than the alternatives - they just move the work from architecture design to inference-time verification, checking and rechecking candidate moves as they go. That shift makes them substantially more expensive in memory and runtime than the graph-based and reinforcement-learning methods, even when solving the identical puzzle.
Translation: making a model sound like it's reasoning through a puzzle costs real compute, and that style of reasoning still can't out-hustle a well-tuned search algorithm on its own turf.