[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-gate-norm-prunes-llm-attention-layers-in-under-a-second":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8700,"gate-norm-prunes-llm-attention-layers-in-under-a-second","Gate-Norm Prunes LLM Attention Layers in Under a Second","A weight-only pruning method removes redundant attention layers from LLaMA models with no training data, lifting throughput up to 1.30x.","A new pruning method shows much of an LLM's self-attention machinery can be deleted without retraining.\n\nA paper titled 'Data-Free Pruning of Self-Attention Layers in LLMs' (arXiv:2512.20636, posted September 30, 2026) introduces Gate-Norm, a method that ranks and removes self-attention sublayers that contribute little to a model's output. The researchers call this the Attention Suppression Hypothesis: during training, some deep attention layers apparently learn to mute themselves, leaving the residual stream and the MLP layers to carry the representation. Gate-Norm scores each layer by query-key coupling and cuts the weakest ones, with no calibration data, no forward passes, no fine-tuning, and no specialized kernels. On 40-layer, 13-billion-parameter LLaMA models, it prunes 8 to 16 attention sublayers in under a second.\n\nThat speed matters as much as the result. Existing pruning techniques usually require calibration datasets and multiple forward passes to decide what to cut, which costs time and infrastructure most teams would rather not spend on compression. Gate-Norm reportedly matches the accuracy of those data-driven methods, keeping zero-shot accuracy within 1.5 percentage points across benchmarks including BoolQ, HellaSwag, and ARC, while scoring layers about 1,000 times faster and delivering up to 1.30x higher inference throughput.\n\nIt is a reminder that a chunk of what looks load-bearing in these architectures may just be dead weight the training process never got around to trimming.","[\"llm-pruning\",\"model-compression\",\"attention-mechanism\",\"arxiv-research\"]","2026-09-30T04:00:00.000Z","2026-09-30T20:11:19.089Z","2026-09-30T20:11:24.377Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Add explicit source attribution — name the paper (\"Data-Free Pruning of Self-Attention Layers in LLMs,\" arXiv:2512.20636), since the draft only says 'researchers describe' without citing the paper, arXiv ID, or date needed to verify the claims.","resolved","ai",[32,33,34,35],"llm-pruning","model-compression","attention-mechanism","arxiv-research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2512.20636",0,{"sections":42},[43,46,50,54,59,64,68,73,78,82,87,92,97,102],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5184,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",791,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",417,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",155,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]