Evolution strategies just got a tune-up, and the fix comes from finding a flaw in the shortcut that made them fast enough for large language models.
Earlier this year, a technique called EGGROLL made it feasible to train huge models with evolution strategies by using low-rank, often rank-one, random perturbations instead of full dense noise. That shortcut is what makes ES scale, but the new paper shows it also warps the underlying math: the rank-one trick can turn the algorithm's effective update into something that no longer behaves like a clean gradient, occasionally flipping the stability of an optimum it should be converging toward. The researchers also found the extra noise from using rank-one instead of dense perturbations shrinks fast as models get wider, down to just 0.098% at a width of 4096. From that diagnosis they built LOO-ROLL, an estimator that gets the same signal from one evaluation per direction instead of two.
That efficiency gain is not just theoretical. Tested across fourteen post-training runs up to 14 billion parameters, LOO-ROLL beat EGGROLL on eleven paired comparisons with no losses, and on math benchmarks the gains were substantial: up to 14.1 points on GSM8K and 12.2 points on MATH-500 at matched compute time.
It is a reminder that the tricks making today's training methods fast are not free lunches, they are trading exactness for speed in ways that only get scrutinized after the fact.