A new transformer design lets individual tokens decide how much extra thinking they need, instead of forcing every word through the same number of reasoning loops.
The architecture, called T-LoopFormer, builds on looped transformers, models that reuse one shared block of parameters across multiple passes to save memory and reason in latent space rather than spelling out every step in text. Earlier looped models applied a fixed number of loops to every token regardless of difficulty, which wasted compute on easy tokens and possibly shortchanged hard ones. T-LoopFormer adds a token-choice router that reads each token's hidden state and decides on the spot how many extra iterations it gets. It also gives each recursion loop its own key-value cache, so tokens at different depths only attend to matching cached states, which the researchers say speeds up autoregressive decoding. On the paper's own language modeling and zero-shot reasoning benchmarks, T-LoopFormer matched or beat prior looped transformers while posting the lowest decoding latency of the models tested.
Fixed-depth reasoning has been one of the quiet costs of latent reasoning models: they save tokens by thinking in hidden states, but they still burn a uniform amount of compute per token whether or not that token needs it. Dynamic per-token depth is essentially mixture-of-experts logic applied to time instead of parameters, routing effort where it's actually needed rather than spreading it evenly.
The results come from the authors' own benchmark suite, so whether the router's depth calls generalize to messier, real-world prompts - and whether the latency gains survive on hardware the authors didn't test - is still an open question.