Turns out figuring out what comes after Wednesday inside a language model looks like solving a two-step problem: work out one relationship, then fold in the rest.
A preprint posted September 30, 2026 to arXiv - arXiv:2609.35970, 'Causal and Interpretable Structures in LLM Compositional Tasks,' not yet peer-reviewed - examined how Llama, Qwen, Gemma, and Mistral models internally solve prompts that require reasoning about cyclic sequences: months, hours, weekdays, and musical notes. The researchers tracked activations across layers and found a consistent pattern across all four model families. Middle layers first encode the relationship between two of the three relevant tokens. Later layers then fold in the third token to build a full three-way representation that actually drives the next-token prediction. The team also found other token relationships that showed up geometrically inside the models but had zero causal effect on the output - structure that exists but does nothing. When the researchers restricted models to rely only on the causally relevant representations, prediction accuracy on these tasks improved.
That's a sharper answer than most interpretability work gives to a basic question: does a model actually compute a relationship, or just pattern-match to something plausible? Finding the same staged, two-then-three-token structure across four unrelated model families suggests transformers may converge on a shared strategy for combining relational information, rather than each just improvising its own.
Worth noting: this is four models tested on toy cyclic sequences, not the messier, many-step reasoning behind real code or math, and the paper hasn't been peer-reviewed - so treat the tidy layered story as a lead, not a settled fact.