Plain gradient descent is bad at learning rare words, and a new minimal model shows exactly why.
Researchers trying to explain why adaptive optimizers beat vanilla stochastic gradient descent on language models stripped away the usual complexity: no transformer layers, no attention, no sequence structure. What is left is a bare softmax unigram model trained on heavy-tailed data, where a handful of common tokens appear constantly and most tokens appear rarely. In that stripped-down setting, the gap between optimizers still shows up. The researchers prove why: gradient descent updates a word's logit in proportion to how often that word appears, so rare-token logits barely move, while SignGD only cares about the direction of the gradient, not its size, so rare and common tokens get updates on a comparable scale. They back this with convergence rate bounds, upper and lower bounds showing gradient descent's slowdown on rare tokens, and upper bounds showing SignGD escapes it, though noise in the stochastic version can hide the advantage unless batch size grows.
This matters because natural language is heavy-tailed by nature: a small set of common words dominates and a long tail of rare words shows up occasionally. That is the exact condition under which adaptive methods have quietly outperformed SGD-style training for years, usually explained with hand-wavy appeals to loss landscape geometry. This paper swaps the hand-waving for a provable mechanism, in the simplest model where the effect still survives.
It is a unigram model, not a full language model, so treat it as a clean explanation of one ingredient in a messier dish, not proof that sign-based updates are a free lunch at full transformer scale.