A new analysis says grokking, the weird delay between a model memorizing its training data and actually generalizing, isn't guaranteed by patience alone.
Researchers built a framework around mode connectivity, mapping the geometry of the low-loss regions a network settles into during training and validation. They reran the standard modular-arithmetic setup that first produced grokking in transformers, but split the data in a way that preserves the task's underlying symmetry. Validation accuracy never recovered. That stable anti-grokking case contradicts several popular explanations that tie grokking to particular hyperparameters or training recipes.
The real driver, the researchers argue, is whether the low-loss regions carved out by the training split and the validation split line up geometrically. When they're misaligned, you get the slow-burn generalization everyone calls grokking. When they're aligned, tuning learning rate or weight decay won't manufacture grokking: the model either trains cleanly or it doesn't.
That's a useful reality check for a field that has spent several years treating grokking as a tunable party trick instead of a geometric accident.