A new paper maps the internal circuitry that decides, within a single token, what language a multilingual model will answer in.
Researchers analyzed six models across four families - GPT-2, BLOOM-560M, Pythia-1B and 2.8B, and Qwen2.5-1.5B Base and Instruct. Using edge attribution patching combined with exact activation-patching verification, they traced which attention heads and layers broadcast the first-token language decision. Broadcasting hubs sit deep or mid-to-deep in most of the networks, though the evidence is strongest for Pythia-2.8B and BLOOM-560M. Scaling Pythia from 1B to 2.8B parameters pulled in more participating nodes while keeping a similar verified-edge budget, producing a sparser circuit; Qwen's base and instruct versions share 84.7% circuit overlap, including a hub at layer 27, meaning instruction tuning barely touches the machinery that sets response language.
This adds a concrete data point to the broader push to open up language models node by node, alongside prior work on induction heads and refusal circuits. Pinning down where language identity gets locked in could help developers debug why a multilingual model suddenly answers in the wrong language, rather than just prompting around the problem.
The paper's own numbers carry a built-in caution: the cheap gradient-based scores used to shortlist candidate circuits correlated only weakly with the slower, exact-patching results. The technique that makes this kind of interpretability work scalable is also the one you can't fully trust without the expensive check.