A new arXiv paper explains why small neural networks keep reaching for the same trick when they learn modular addition.
Researchers studied two-layer networks trained on modular addition, a toy task where models often discover Fourier-structured representations that generalize exactly. Prior work had spotted these Fourier circuits but not explained why gradient descent finds them. The new paper introduces "probability signatures," a way of tracking the leading gradient interactions through conditional statistics of the training data. For modular addition, those signatures turn out to be cyclic shift operators that the discrete Fourier transform neatly untangles into separate, roughly independent frequency modes, explaining the sparsity, frequency matching, and phase alignment seen in trained networks. The same framework also explains an odd observation: mislabeled training examples can produce faster early loss drops than clean ones, even though they carry no real generalization signal, because label noise increases coincidental agreement between examples that reinforces shared directions early in training.
This matters because interpretability research keeps finding tidy internal structures in neural networks without a solid account of why training produces them. A mechanistic explanation, grounded in the data's statistics rather than post-hoc pattern spotting, is a sturdier foundation for understanding how models generalize and when noisy data might fool early training metrics.
The authors also tested the method on XOR and found the predicted frequency pattern there too, suggesting the approach generalizes beyond modular addition. Still, this is toy-task interpretability. Whether these clean, diagonalizable dynamics say anything about the tangled optimization landscape of a production-scale language model is, for now, an open question.