AI/ ai · deep learning · optimizers · ai research

New Proof Shows Muon Optimizer Reaches Target Loss

A new proof shows the Muon optimizer, popular in large AI training runs, reliably hits a specific loss target rather than just stalling at a plateau.

A new paper gives one of the first real proofs that Muon, the momentum-based optimizer spreading through large AI training runs, actually drives a network to a specific target loss, not just some vague stopping point.

Muon's speed comes from a step called Newton-Schulz orthogonalization, which in practice runs for only five tuned iterations rather than running to full precision. Prior analyses either assumed that step was exact or used the classic version of the math, and could only prove the network reaches a flat spot, not a usable loss level. The new work analyzes the real five-step version, including the momentum buildup that happens before it, and proves that for a sufficiently wide two-layer ReLU network, Muon reaches any target training loss with high probability. The time needed scales with the inverse square root of that target loss, divided by how much momentum budget is left, and the required network width does not depend on the target or the momentum setting. The authors back this with 30 runs across six widths and five random initializations, all of which hit the target loss.

Muon has already been adopted in several open-source training stacks as a faster alternative to Adam, but its five-step shortcut was something people trusted because it worked in benchmarks, not something anyone had proven. This result gives that shortcut a mathematical basis, at least in a simplified setting, which matters for anyone deciding whether to bet a training run on it instead of a better-understood optimizer.

Two-layer networks with fixed random output weights are a long way from the billion-parameter transformers Muon is actually tuned for in production, so the gap between proof and practice has narrowed, not closed.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →