[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-study-finds-training-order-effects-depend-on-learning-rate-decay":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},9454,"study-finds-training-order-effects-depend-on-learning-rate-decay","Study Finds Training Order Effects Depend on Learning Rate Decay","A new study finds that training-order effects on language models vanish under a convention-agnostic score, exposing a blind spot in how benchmarks are built.","Order your training data differently and a language model commits to something different, but only if you also decay the learning rate. A new analysis finds that effect, and shows it vanishes completely once you use a scoring method that doesn't play favorites between conventions.\n\nResearchers trained models on ten different orderings of the same corpus, where every problem could be written under two equally correct but incompatible conventions. Under a constant learning rate, shuffling the order moved which convention the model favored, with a measured gap of 11.63 across the ten runs. Switch to the cosine-decay schedule used in nearly every published training run, and the same ten orderings collapsed into just two sharply separated outcomes, a difference the paper reports at 12.29 sigma. The data's arrangement and the schedule's shape multiply together to produce the effect; neither one alone explains it.\n\nAcross twelve runs, combined accuracy on both conventions stayed flat, within 9.7% of constant, even as the model's preference for one convention over the other swung from 4% to 87%. The model was not getting smarter or dumber. It was committing to a formatting habit that a standard accuracy benchmark cannot see, because exact-match scoring only checks whether an answer is right, not which convention produced it.\n\nSimply naming the convention in the prompt erased most of that swing, recovering 87.5% of the best possible combined score. That is a cheap fix for what sounds like a scary instability, and a reminder that reported training randomness in language models may be measurement blindness dressed up as chaos.","[\"machine-learning\",\"training-data\",\"benchmarks\",\"language-models\"]","2026-10-02T04:00:00.000Z","2026-10-02T19:49:51.599Z","2026-10-02T19:49:57.397Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The 12.29-sigma figure is misattributed to the constant-learning-rate run — per the source, that run's own measured gap is 11.63, and 12.29 is the headline figure quoted from the decayed-schedule family; fix the attribution so the constant-rate number matches the source.","resolved","ai",[32,33,34,35],"machine-learning","training-data","benchmarks","language-models",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00234",0,{"sections":42},[43,46,50,55,60,65,70,75,80,85,90,95,100,105],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5764,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",831,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",437,"2026-10-01T18:10:00.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",198,"2026-10-01T17:38:48.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",168,"2026-10-01T18:35:55.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]