AI/ ai · llm-training · optimizers · machine-learning

New Optimizer Trims Memory for AI Model Vocabulary Layers

A new optimizer variant called MuonIO cuts memory and compute for the embedding and output layers that dominate large language model training costs.

Researchers have extended the Muon optimizer to cover the two layers it used to skip.

Muon speeds up neural network training by treating weight updates as a norm-constrained problem rather than a plain gradient step. Standard implementations, though, fall back to the older AdamW optimizer for two specific layers: the embedding table, which turns tokens into vectors, and the language model head, which turns vectors back into predictions over the vocabulary. A new method called MuonIO applies Muon-style logic to both, using a different operator norm for each - the 2-to-infinity norm for the output head and the 1-to-2 norm for the embedding table - chosen to match how each layer actually processes information. In practice, the update looks like row normalization for the head and column normalization for the embedding table. Tested on a 1 billion-parameter LLaMA-style model trained on the C4 dataset, it cut optimizer memory for those two layers in half, reduced their update compute by about 46 percent, and still improved validation perplexity.

That matters because embedding tables and output heads scale with vocabulary size, not just model size, so they eat a disproportionate share of memory in models with large or multilingual vocabularies. Shaving 50 percent off that overhead is a real systems win even though it touches no attention block and invents no new architecture.

It is the kind of unglamorous fix that will not headline a product launch, but it is exactly the sort of detail that shows up later in somebody's training bill.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →