AI/ llm pretraining · data selection · optimizers · ai research

Study Finds Optimizer Choice Determines Data Selection Gains

A new theoretical study shows that AdamW-style optimizers squeeze far more value out of smart data selection during LLM pretraining than Lion or Muon do.

Not every optimizer plays nice with smart data selection, according to a new theoretical study on LLM pretraining.

The paper, called OptiSelect, formalizes an "optimizer-aware" approach to online data selection, the practice of training only on the most valuable examples in a batch rather than all of them. The researchers prove that optimizers using sign-based or polar-tangential preconditioning, the techniques behind Lion and Muon, suffer what they call a "discriminability collapse" that caps how much a smart selection scheme can actually help. Diagonal-adaptive optimizers like AdamW and Sophia, by contrast, have strictly better theoretical ceilings. Pretraining runs on 124M and 720M parameter models backed up the math: AdamW's scoring geometry stayed the strongest option for ranking data, even in setups that used Muon as the optimizer.

That is a useful, if unglamorous, finding for anyone optimizing training budgets. Teams often treat optimizer choice and data curation as separate knobs to tune independently, but this work suggests they interact, and pairing a newer optimizer like Muon with a selection pipeline could quietly waste the pipeline's benefit. The study also derives an optimal oversampling ratio for candidate selection and shows the approach still holds up when training data has been rephrased, a common step in modern data pipelines.

In other words, before chasing the latest optimizer trend, check whether your data selection logic actually benefits from it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →