[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-standard-network-initialization-destabilizes-derivative-losses":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},6319,"standard-network-initialization-destabilizes-derivative-losses","Standard Network Initialization Destabilizes Derivative Losses","A new analysis shows the initialization scheme that stabilizes wide networks lets derivative variance grow with depth, undermining derivative-based losses.","A new paper finds a crack in the initialization trick that makes deep neural networks trainable in the first place.\n\nThat trick is called critical, or 'edge of chaos,' initialization. It keeps signals from exploding or vanishing as they pass through many layers, by tuning weight variance so a small nudge to the input produces a proportional nudge to the output, no matter how deep the network gets. The researchers studied simple scalar-input networks and worked out the math for how derivatives beyond the first behave under this scheme. They found the first derivative stays stable with depth, as expected, but the second derivative's variance grows linearly with depth whenever the activation function curves, which most do. For a residual network variant that scales each layer's contribution by one over the square root of depth, they proved every derivative order stays bounded instead.\n\nThat gap matters because a growing set of training methods - physics-informed neural networks, score-based generative models, and any loss that penalizes a network's derivatives directly - lean on those higher derivatives being well-behaved. If second-derivative variance balloons with network depth, gradients for these losses could become noisy or unreliable exactly where you'd want a deep network to help most.\n\nThis is initialization theory, not a benchmark result, so no model got better or worse today. But it is a reminder that 'critical' was never a general-purpose guarantee - it was tuned for first derivatives, and anything reaching for the second one should check the math before scaling up.","[\"neural networks\",\"initialization\",\"deep learning theory\",\"physics-informed ml\"]","2026-09-11T04:00:00.000Z","2026-09-11T06:15:38.314Z","2026-09-11T06:15:50.215Z","published",null,[],"ai",[26,27,28,29],"neural networks","initialization","deep learning theory","physics-informed ml",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.09244",0,{"sections":36},[37,40,44,48,53,58,63,66,71,75,80,85,90,95],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",3521,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",637,{"name":45,"slug":46,"count":47,"latest_published_at":18},"Policy","policy",338,{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":57},"Hardware","hardware",153,"2026-09-09T15:12:32.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":64,"slug":65,"count":61,"latest_published_at":18},"Science","science",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":18},"Dev Tools","dev-tools",70,{"name":76,"slug":77,"count":78,"latest_published_at":79},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]