Large language models' reasoning accuracy collapses on deep problems unless they are allowed to write and run code instead of just thinking in words.
Researchers built a benchmark around Boolean circuits over GF(2), a mathematical structure that can represent any computable function. The design strips out two things that make typical reasoning benchmarks hard to trust: the chance a model memorized the answer, and the ambiguity in how a problem can be broken into steps. Models solved these circuits step by step while researchers controlled circuit depth directly. When reasoning purely in natural language, next-step accuracy fell apart as depth increased, a pattern that held for both small and frontier models.
That matters because out-of-distribution reasoning, the ability to recombine learned rules to solve problems a model has not seen before, is the capability underpinning claims about general intelligence in AI systems. The benchmark suggests that capability is shakier than headline scores imply. When models were instead allowed to synthesize and execute code to solve the same circuits, the depth-related collapse did not happen, in either the small or frontier models tested.
The fix for brittle reasoning here is not more training data. It is a calculator. That is a humbling finding for systems marketed as reasoning engines.