AI/ ai · machine-learning · optimization · research

A Fixed Version of Muon for Convolutions Works No Better

A team built a mathematically correct Muon optimizer for convolutions, and it performed no better than the shortcut version everyone already uses.

Muon, a popular optimizer for training neural networks, is built on math that doesn't actually apply to convolutional layers - and fixing that math didn't make it better.

Muon works by computing a polar factor for matrix-shaped weight updates. Convolutional kernels are stored as four-dimensional tensors, so standard implementations just reshape them into matrices to use the same trick, even though that breaks the theoretical reasoning behind why Muon works. A group of researchers built an alternative called Conv-NS that does the polar-factor math directly in the actual convolutional geometry, preserving the kernel's structure instead of flattening it. They tested both versions on CIFAR-10 and ImageNet image classification and found the theoretically correct Conv-NS trained about as efficiently and landed on comparable accuracy to the reshape-based hack.

That's a strange result: the "wrong" method and the "right" method tie. The researchers' working theory is that forcing exact orthogonalization onto convolutional updates - the mathematically pure move - may actually overconstrain the optimizer rather than help it. It's a reminder that in deep learning, an optimizer's success is often empirical first, theory a distant second.

Call it the Muon paradox: the shortcut survives the theory police and comes out even.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →