New research says AI reasoning models may be weighting their own thought process wrong.
A paper posted to arXiv (arXiv:2609.13747, posted September 2026) examines "continuous reasoning," a technique where a language model holds multiple lines of reasoning in a single latent state instead of writing each one out as text. The researchers tested a common assumption: that a model should mostly preserve its most recent reasoning step, since older steps compete for limited hidden space. They found the opposite holds when later computation draws on several earlier ideas at once. In that case, weighting all reached ideas equally lets a model guide its attention correctly using less hidden width than an approach that only keeps the newest states. Tests on two-layer and GPT-2 Transformers confirmed the pattern, and showed that when weights are uneven, the least-weighted ideas are the first to degrade.
This is a narrow, technical result, but it points at something practical. Frontier reasoning models today spend enormous compute writing out and re-reading chains of thought as text. Continuous reasoning tries to do that work inside the model's hidden layers instead, which is cheaper if researchers can get the bookkeeping right. This paper offers a concrete default for that bookkeeping: keep every live idea equally weighted, and reset to equal weights as computation proceeds.
It won't show up in a shipped product this year, but it's the unglamorous math that will decide whether reasoning models get meaningfully cheaper to run, or just get bigger.