AI/ optimization · llm-training · ai-research

New Optimizer Framework Cuts LLM Training Token Costs by 9.6%

A new paper unifies Muon-style optimizers under one framework and shows a variant called TensorChain trains Qwen3 models using meaningfully fewer tokens.

A new optimizer called TensorChain trains language models to the same accuracy using fewer tokens than Muon, the optimizer it is built on.

Researchers built a framework called chained linear minimization oracles, or chained LMOs, to explain a growing family of optimizers descended from Muon, which works by normalizing matrices during training. They found that many of these chained optimizers technically violate the framework's own rules and can fail to converge on simple, well-behaved problems - yet they still perform well in practice. To explain that contradiction, the team turned to linear associative memory theory and showed chaining helps specifically when a model's embeddings are anisotropic, meaning they are stretched unevenly across dimensions. Building on that insight, they designed TensorChain, which stacks weight matrices from multiple layers into a single 3d tensor and normalizes it across all three axes at once.

Training cost scales with tokens processed, so efficiency gains here translate directly into compute and dollars saved. In pretraining runs on Qwen3 0.6B and 1.7B, TensorChain hit the same validation loss as Muon using 9.6% fewer tokens on average, and it beat every other chained-optimizer baseline tested.

Optimizer papers have a habit of posting strong numbers at sub-2-billion-parameter scale that quietly evaporate once someone tries them at 70 billion parameters and up - so treat this as a promising lead, not a settled upgrade.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →