[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-loss-landscape-geometry-explains-why-llm-training-needs-warmup":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},9172,"loss-landscape-geometry-explains-why-llm-training-needs-warmup","Loss Landscape Geometry Explains Why LLM Training Needs Warmup","Researchers map how pretraining loss landscapes change over time, turning the finding into concrete rules for learning rate warmup and batch size scheduling.","A new paper argues the geometry of a language model's loss landscape - not trial and error - should set the pace for learning rate warmup and batch size.\n\nResearchers tracked how the local shape of the loss landscape changes as a language model pretrains, splitting the process into two phases. In Phase I, the landscape starts out sharp, which makes training unstable and causes loss plateaus if the learning rate is too high. That sharpness eases early in training, which is why warmup helps - and the paper's theory suggests a bigger peak learning rate needs a proportionally longer warmup to match. In Phase II, the landscape's shape is driven by gradient noise: smaller batches add noise that widens the loss basin, while larger batches reduce noise and deepen it. That depth-flatness trade-off motivates a batch size schedule that starts small and grows late in training.\n\nHyperparameter tuning for pretraining runs is expensive and still largely guesswork, by the paper's own framing - there's no principled theory connecting warmup length, learning rate, and batch size to training stability. This work offers one: a geometric explanation for why warmup works and a concrete recipe - start with a small batch, grow it later - instead of another round of grid search.\n\nThe ideas are clean, but they come from theory and the paper's own experiments. Whether a rising batch size schedule survives contact with a trillion-parameter training run, with all its infrastructure quirks, is a separate question nobody has answered yet.","[\"ai research\",\"llm pretraining\",\"hyperparameter tuning\"]","2026-10-01T04:00:00.000Z","2026-10-01T23:52:30.081Z","2026-10-01T23:52:32.221Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Remove or substantiate the unsupported claim that hyperparameter tuning burns 'compute budgets measured in millions of dollars per run' — this specific cost figure appears nowhere in the source material and reads as an invented statistic.","resolved","ai",[32,33,34],"ai research","llm pretraining","hyperparameter tuning",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39767",0,{"sections":41},[42,45,49,53,58,63,67,72,77,81,86,91,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5598,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",815,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",430,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":18},"Science","science",163,{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":78,"slug":79,"count":75,"latest_published_at":80},"Software","software","2026-09-30T21:41:11.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]