[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-new-llm-optimizer-claims-wins-over-adamw-and-muon":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8604,"a-new-llm-optimizer-claims-wins-over-adamw-and-muon","A New LLM Optimizer Claims Wins Over AdamW and Muon","Researchers unveil NormPre, an optimizer that beats AdamW, Muon, and MANO in pretraining tests, though the paper skips the actual numbers.","A new optimizer claims it beats the standard toolkit for training large language models, but the paper never says by how much.\n\nResearchers built NormPre by rethinking Muon, an optimizer that treats a weight update as one combined signal mixing each parameter's overall scale with how parameters interact geometrically. NormPre splits that signal into two steps: first it normalizes marginal scale using diagonal-Gram information, then it applies spectral preconditioning to the directional interaction geometry. Two variants came out of this: NormPre-G, which uses Newton-Schulz iterations for global preconditioning, and NormPre-L, which uses randomized sketching to approximate the leading interaction eigenspace for cheaper, localized preconditioning. The team ran pretraining experiments on GPT-2 Small, LLaMA, and Qwen3 and reports both variants \"consistently outperform\" AdamW, Muon, and MANO under matched training budgets, plus O(T^-1\u002F2) convergence guarantees for simplified versions.\n\nThat's the entire comparison the abstract offers. It does not report loss values, perplexity, downstream eval scores, or any percentage gain, so there's no way to judge whether beating AdamW and Muon here is a rounding error or a real jump. Optimizer papers live and die on exactly those numbers, and matrix-based methods like Muon have already drawn scrutiny over whether their gains survive outside curated benchmarks.\n\nThe code is open-source on GitHub, so other labs can run their own head-to-head tests instead of taking the abstract's word for it. Until someone does, \"outperform\" here is a claim, not a measurement.","[\"llm-training\",\"optimizers\",\"open-source\",\"ai-research\"]","2026-09-30T04:00:00.000Z","2026-09-30T14:18:54.557Z","2026-09-30T14:19:00.610Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The headline, dek, and body assert NormPre 'outperforms' AdamW\u002FMuon\u002FMANO without ever stating what metric was measured (loss? perplexity? downstream eval?) or citing the magnitude of the improvement — add the specific metric and comparison figures from the paper, or explicitly note the abstract doesn't disclose numeric results if that's the case.","resolved","ai",[32,33,34,35],"llm-training","optimizers","open-source","ai-research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.36692",0,{"sections":42},[43,46,50,54,59,64,69,74,79,83,88,93,98,103],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5135,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",788,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",417,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]