[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-math-settles-a-long-debate-over-momentum-in-gradient-descent":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},9311,"math-settles-a-long-debate-over-momentum-in-gradient-descent","Math Settles a Long Debate Over Momentum in Gradient Descent","A new theoretical analysis shows SGD with momentum stays just as stable and generalizable as plain SGD, challenging a common training folk wisdom.","A new theoretical study throws cold water on a popular piece of machine learning lore: that momentum makes training faster but secretly makes models worse at handling new data.\n\nResearchers built a generalization analysis of stochastic gradient descent with momentum (SGDM) using a framework called algorithmic stability, which measures how much a model's output changes if you swap one training example for another. The analysis covers both Polyak's and Nesterov's momentum, the two schemes used across nearly all modern optimizers. The team proved stability bounds that hold for any momentum value between 0 and 1, for smooth and convex problems, and without assuming the loss function is Lipschitz continuous, a technical condition most prior work required. Combining these with new bounds on optimization error, they derived what they call optimal excess population risk bounds, essentially a formal guarantee on how well a trained model should perform on data it has never seen.\n\nThat result matters because the tradeoff between training speed and generalization has mostly been managed by intuition and hyperparameter sweeps, not proof. Momentum has been standard practice since the 1980s precisely because it trains faster, and plenty of practitioners have quietly worried it was borrowing that speed from generalization performance. This paper gives a mathematical reason to stop worrying, at least in the convex setting where the guarantees hold.\n\nThe obvious caveat: smooth and convex problems describe textbook optimization, not the sprawling non-convex loss surfaces of a modern transformer. Expect this result to shape how optimization theorists frame the debate long before it changes what a deep learning engineer actually does on Monday morning.","[\"machine learning\",\"optimization theory\",\"sgd\",\"research\"]","2026-10-01T04:00:00.000Z","2026-10-02T08:45:18.195Z","2026-10-02T08:45:18.881Z","published",null,[],"ai",[26,27,28,29],"machine learning","optimization theory","sgd","research",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.28517",0,{"sections":36},[37,40,44,48,53,58,62,67,72,76,81,86,91,96],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",5690,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",820,{"name":45,"slug":46,"count":47,"latest_published_at":18},"Policy","policy",430,{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",308,"2026-10-01T12:30:00.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":57},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",165,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",150,"2026-10-01T11:59:27.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":73,"slug":74,"count":70,"latest_published_at":75},"Software","software","2026-09-30T21:41:11.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]