Order your training data differently and a language model commits to something different, but only if you also decay the learning rate. A new analysis finds that effect, and shows it vanishes completely once you use a scoring method that doesn't play favorites between conventions.
Researchers trained models on ten different orderings of the same corpus, where every problem could be written under two equally correct but incompatible conventions. Under a constant learning rate, shuffling the order moved which convention the model favored, with a measured gap of 11.63 across the ten runs. Switch to the cosine-decay schedule used in nearly every published training run, and the same ten orderings collapsed into just two sharply separated outcomes, a difference the paper reports at 12.29 sigma. The data's arrangement and the schedule's shape multiply together to produce the effect; neither one alone explains it.
Across twelve runs, combined accuracy on both conventions stayed flat, within 9.7% of constant, even as the model's preference for one convention over the other swung from 4% to 87%. The model was not getting smarter or dumber. It was committing to a formatting habit that a standard accuracy benchmark cannot see, because exact-match scoring only checks whether an answer is right, not which convention produced it.
Simply naming the convention in the prompt erased most of that swing, recovering 87.5% of the best possible combined score. That is a cheap fix for what sounds like a scary instability, and a reminder that reported training randomness in language models may be measurement blindness dressed up as chaos.