A team of mathematicians has worked out a rigorous theory for one of deep learning's stranger habits: training that bounces instead of smoothly descending.
When a neural network trains with a learning rate large enough, its loss doesn't fall smoothly. It oscillates, a pattern known as the edge of stability, and researchers have long noticed this jittery process often produces models that generalize better than careful, slow training. A new paper introduces mean-fluctuation dynamics, a continuous-time model that tracks both the averaged training path and the size of its oscillations as linked variables. The authors derive the model rigorously from gradient descent in a simplified sharp-valley setup, map its stable resting points, and extend the analysis to wide two-layer neural networks, proving the math holds up and identifying conditions under which training provably converges.
Edge-of-stability training already happens by default across much of deep learning. It's usually a side effect of picking a learning rate large enough to train fast, not a deliberate strategy. A provable model for why that wobbly process can still converge, and converge well, gives researchers a mathematical foothold for learning-rate choices that have mostly been justified by trial and error and GPU budgets.
The strongest guarantees here still live in two-layer-network territory, so whether this framework scales to the sprawling architectures behind today's large language models remains open work.