A new pruning method shows much of an LLM's self-attention machinery can be deleted without retraining.
A paper titled 'Data-Free Pruning of Self-Attention Layers in LLMs' (arXiv:2512.20636, posted September 30, 2026) introduces Gate-Norm, a method that ranks and removes self-attention sublayers that contribute little to a model's output. The researchers call this the Attention Suppression Hypothesis: during training, some deep attention layers apparently learn to mute themselves, leaving the residual stream and the MLP layers to carry the representation. Gate-Norm scores each layer by query-key coupling and cuts the weakest ones, with no calibration data, no forward passes, no fine-tuning, and no specialized kernels. On 40-layer, 13-billion-parameter LLaMA models, it prunes 8 to 16 attention sublayers in under a second.
That speed matters as much as the result. Existing pruning techniques usually require calibration datasets and multiple forward passes to decide what to cut, which costs time and infrastructure most teams would rather not spend on compression. Gate-Norm reportedly matches the accuracy of those data-driven methods, keeping zero-shot accuracy within 1.5 percentage points across benchmarks including BoolQ, HellaSwag, and ARC, while scoring layers about 1,000 times faster and delivering up to 1.30x higher inference throughput.
It is a reminder that a chunk of what looks load-bearing in these architectures may just be dead weight the training process never got around to trimming.