A new study finds that an old regularization trick quietly decides how extreme a transformer's strangest neurons get.
Researchers tracked the life cycle of massive activations, residual-stream values far bigger than normal activations that are linked to attention sinks. They found that which channels carry the sink varies across random seeds but locks in early in each training run. As training continues, neighboring channels fade out, and the sink concentrates onto a small set of redundant carriers. Through controlled training interventions, the team showed that weight decay drives the rise and fall: removing it near the peak lets the scale keep climbing, while keeping it causes a decline even at a constant learning rate.
This matters because gradient flow through the network tracks the collective size of these activations, not any single channel, which affects how models learn and how they hold up under pruning or quantization. The finding hands engineers a concrete lever, the decay coefficient, for shaping a model's internal behavior without touching its architecture.
It is a useful reminder that today's largest models still run on tuning knobs borrowed from far smaller networks decades ago.