[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-proof-shows-why-small-neural-nets-default-to-simple-fits":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},4882,"new-proof-shows-why-small-neural-nets-default-to-simple-fits","New Proof Shows Why Small Neural Nets Default to Simple Fits","A new theoretical paper proves why gradient descent on quadratic neural networks converges to the simplest solution that fits the data.","A new arXiv paper pins down the exact math behind why gradient descent training on a class of simple neural networks reliably lands on the simplest possible answer, not just any answer that fits.\n\nThe paper studies \"positive quadratic networks,\" models of the form f(x) = x^T UU^T x, where the matrix U is redundant: rotating it doesn't change the network's output, so many different U's describe the same function. The authors show that training these networks with gradient descent is mathematically identical to a cleaner process on the space of valid solutions once that redundancy is stripped out. Using this framing, they work out the curvature of the loss landscape at solutions that perfectly fit noisy random data, prove bounds on how fast training converges, and build a better way to initialize the network's weights. They also show that under very small starting weights, training converges to the solution with the smallest possible \"trace,\" roughly the least complex fit, and they back this with numerical experiments.\n\nThis is implicit bias research: figuring out why gradient descent, with no explicit rule telling it to prefer simple models, tends to pick simple ones anyway. That question sits underneath a lot of why deep learning generalizes well despite having far more parameters than training examples, and quadratic networks are a tractable stand-in for studying it rigorously. The same math also applies to matrix sensing and recovery problems, the kind of setup used in compressed sensing and some imaging applications.\n\nIt is proof-heavy and synthetic-data-tested rather than something you will see in a product changelog, but it is the sort of groundwork that eventually explains why the black box works at all.","[\"ai\",\"machine-learning\",\"research\",\"optimization\"]","2026-07-30T04:00:00.000Z","2026-08-14T04:56:26.222Z","2026-08-14T04:56:38.006Z","published",null,[],"ai",[24,26,27,28],"machine-learning","research","optimization",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.25624",0,{"sections":35},[36,40,44,49,54,59,64,69,74,79,84,89,94,99],{"name":37,"slug":24,"count":38,"latest_published_at":39},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":41,"slug":42,"count":43,"latest_published_at":39},"Security","security",435,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]