AI/ muon optimizer · optimization theory · deep learning training · ai research

Researchers Finally Prove Why the Muon Optimizer Works

A new proof shows the popular Muon optimizer actually converges, and extends the guarantee to a Nesterov based variant called Muesterov.

A new paper offers the first real convergence proof for Muon, the training algorithm that plenty of AI labs already use without anyone being able to show it actually converges.

The authors swap out the usual matrix sign function inside Muon's Newton-Schulz step for a more accurate mathematical stand-in, then use it to prove that gradient norms shrink toward zero as training runs. Under a stronger condition known as the global Polyak-Lojasiewicz property, they show the loss itself falls at a linear rate. The key move: the regularization baked into Muon's Newton-Schulz update creates a bounded preconditioner, which means Muon is secretly a preconditioned version of the classic Polyak heavy-ball method. That reframing let the authors borrow a decades-old style of analysis, Lyapunov functions, instead of inventing new machinery. They then applied the same trick to Nesterov's accelerated gradient method, producing a new variant they call Muesterov, and proved it converges under the same conditions.

This matters because Muon spread through practice before the theory caught up, which is the usual order for optimizers in deep learning. A formal convergence guarantee does not change how Muon performs tomorrow, but it hands researchers building the next optimizer a reusable recipe: precondition, then bolt on heavy-ball or Nesterov momentum, and the proof comes along for free.

The catch is that the evidence behind this is a scalar cross-entropy toy problem and, by the authors' own description, preliminary runs on the small nanoGPT setup. Proving an optimizer works on a whiteboard and proving it works at the scale labs actually train at are still two different exercises.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →