[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-shrinking-ai-models-is-about-wiring-not-weights":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},5621,"shrinking-ai-models-is-about-wiring-not-weights","Shrinking AI Models Is About Wiring, Not Weights","A new study on Pythia language models shows shrinking a big model into a smaller one saves training tokens, but the trick fails at larger scale gaps.","You don't have to train a smaller AI model from scratch: a new paper shows you can wire one together from a bigger sibling's internals, for a lot less than the cost of starting over.\n\nThe researchers tested this on the Pythia model family, converting a 1.4 billion parameter model down to 410 million parameters. They found that the internal representations of large and small models line up well statistically, with a fit score of 0.84 out of a possible 1, but the raw weights do not: directly copying and projecting weights breaks the model's rotary position embeddings, attention heads, and layer norms, and once you correct for the best possible linear transformation, what's left over looks like random noise. So the real value isn't in transplanting weights, it's in using the big model to give the small one a better starting point. The team split that starting point into two adjustable settings, one tuned for best immediate performance and one tuned for best long-run training dynamics, and at 30 million tokens the performance-tuned version beat the leading alternative shrinking method on two separate size reductions.\n\nThat token efficiency is the real story. This approach can hit a given quality bar using up to 18 times fewer training tokens than training from scratch at the low end of the budget range, though the advantage narrows as the training budget grows. It also beats the current standard shrinking pipeline, pruning plus distillation, and does better still when combined with it.\n\nIt's not a universal shrink ray, though: push the donor model to about 5 times the target's size and the correction math starts to overcorrect, so this works best between close relatives in the same model family, not across wildly different sizes.","[\"ai\",\"model-compression\",\"pythia\",\"training-efficiency\"]","2026-08-18T04:00:00.000Z","2026-08-19T04:04:59.479Z","2026-08-19T04:05:11.319Z","published",null,[],"ai",[24,26,27,28],"model-compression","pythia","training-efficiency",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.02829",0,{"sections":35},[36,40,44,49,54,59,64,69,74,78,83,88,93,98],{"name":37,"slug":24,"count":38,"latest_published_at":39},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":41,"slug":42,"count":43,"latest_published_at":39},"Security","security",435,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":18},"Dev Tools","dev-tools",69,{"name":79,"slug":80,"count":81,"latest_published_at":82},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]