Muon, a popular optimizer for training neural networks, is built on math that doesn't actually apply to convolutional layers - and fixing that math didn't make it better.
Muon works by computing a polar factor for matrix-shaped weight updates. Convolutional kernels are stored as four-dimensional tensors, so standard implementations just reshape them into matrices to use the same trick, even though that breaks the theoretical reasoning behind why Muon works. A group of researchers built an alternative called Conv-NS that does the polar-factor math directly in the actual convolutional geometry, preserving the kernel's structure instead of flattening it. They tested both versions on CIFAR-10 and ImageNet image classification and found the theoretically correct Conv-NS trained about as efficiently and landed on comparable accuracy to the reshape-based hack.
That's a strange result: the "wrong" method and the "right" method tie. The researchers' working theory is that forcing exact orthogonalization onto convolutional updates - the mathematically pure move - may actually overconstrain the optimizer rather than help it. It's a reminder that in deep learning, an optimizer's success is often empirical first, theory a distant second.
Call it the Muon paradox: the shortcut survives the theory police and comes out even.