New optimizer math finally has proofs to match the hype.
A new paper works out the theory behind "second-moment stochastic approximation" methods, a category that includes widely used deep-learning optimizers like Adam and its variants, plus the newer Muon optimizer. The researchers frame these methods as solving an optimal preconditioning problem for matrix equations, then build a two-stage convergence analysis. Stage one covers idealized versions that use the exact first and second moments of the underlying random function. Stage two swaps in the estimated moments actually used in practice and, via Dvoretzky's theorem, shows the resulting methods converge almost surely into a neighborhood around the target solution, with the neighborhood's size set by the bias and variance of those estimators. The paper works out concrete convergence bounds for Muon and for a spectral variant of Adam.
That matters because Adam has run deep learning training for roughly a decade on the strength of empirical results, not airtight proofs. Muon, a much newer arrival, has spread through large model training largely on similar faith. This paper gives both a common mathematical home, and a template for judging whichever optimizer shows up next.
Nobody is switching optimizers because of a convergence bound. But papers like this are usually what turns a promising heuristic into the thing every framework ships by default.