Researchers say they've found the internal circuit that makes language models bluff instead of admitting they don't know.
The team studied ten language models, from 3B to 14B parameters, spanning five model families and three benchmarks. Using a technique called causal gating, they isolated a small, consistent set of attention heads and MLP layers, which they call the Commit-Abstain Circuit, that governs whether a model commits to an answer or abstains. Across every model tested, the same pattern showed up: components that push toward committing build confidence in earlier layers, while the components that would trigger abstention only activate later and usually can't override that earlier momentum. In other words, models aren't lying so much as overcommitting before the internal case for uncertainty gets a real vote.
Most hallucination fixes so far work from the outside, flagging or filtering bad answers after they're generated. This traces the problem to its source and shows it can be steered directly: a lightweight policy trained on the circuit's own activations improved commit-or-abstain accuracy by 12.2 points over the model's default judgment, and cut false abstentions by 2.5 times. It also transferred to benchmarks it wasn't trained on and held up on larger 27B-35B models.
It's still a lab result, not a shipped feature. Getting a 14B model to hedge is a different challenge than doing it at frontier scale, where the incentive to sound confident is even stronger.