Reasoning models aren't just thinking longer. They're thinking differently, and a new study shows exactly how.
Researchers ran a needle-in-a-haystack test: bury a bunch of records in a long document and ask a model to count them. Models using chain-of-thought (CoT) reasoning beat non-reasoning models, with the gap widening as the count went up. Digging into why, the researchers found two distinct strategies. Non-thinking models scan broadly, spreading attention across many needles at once. Thinking models do something closer to manual labor: they enumerate needles one by one in their CoT trace, concentrating attention on each in turn and building a more compact internal representation as they go.
The more striking finding is that these models appear to maintain something like an internal counter, updating it as each needle gets retrieved, even when the CoT text never explicitly numbers anything. That's a concrete mechanism behind what's often a black-box claim: that CoT "helps the model think." Here it looks more like CoT gives the model scratch space to track state, one item at a time, which matters for anyone building tools that lean on long-context retrieval, like contract review or log analysis.
Worth noting: the causal evidence for internal counters comes from small, controlled experiments, not the twelve-model comparison itself. Whether this targeted-retrieval behavior holds up on messier, real-world documents is still an open question.