A research team has built a reasoning model with just 27 million parameters that beats billion-parameter language models at logic puzzles.
Researchers describe the Hierarchical Reasoning Model, a recurrent neural network inspired by how the brain splits slow, abstract planning from fast, detailed computation. It uses two interacting modules instead of one monolithic transformer, and skips chain-of-thought prompting entirely. Trained on just 1000 examples with no pretraining, HRM solved complex Sudoku puzzles and found optimal paths through large mazes with near-perfect accuracy. It also beat much larger language models that have far longer context windows on the Abstraction and Reasoning Corpus, a benchmark often cited as a proxy for general intelligence.
Chain-of-thought prompting is the default reasoning trick for today's LLMs, but it is expensive, slow, and falls apart when a task does not decompose cleanly into text steps. HRM suggests a narrower, task-specific architecture can match or beat that approach without mountains of training data or long inference chains, though it is unclear whether a model this specialized generalizes beyond puzzle-style tasks.
Sudoku and mazes are clean, well-defined problems; the real test is whether this hierarchical approach holds up on messier, real-world reasoning tasks where the optimal path is not so obvious.