[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-new-way-to-measure-what-language-models-actually-compute":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},6546,"a-new-way-to-measure-what-language-models-actually-compute","A New Way to Measure What Language Models Actually Compute","A new method shows the popular variance-based way of sizing up a language model's internals often flags directions that barely affect its output.","A new study says the standard trick for sizing up what's happening inside a language model has been measuring the wrong thing.\n\nResearchers tested six models across three families, ranging from 70 million to 7 billion parameters, using a technique called task-weighted charts. Instead of fitting low-dimensional coordinates to raw variance in a model's activations, the method fits them to a specific function, such as next-token prediction, under that function's own metric. They found next-token prediction needs 70 to 90 percent of the residual stream's width just to stay within 5 percent of the model's original perplexity, and nearly all of that width goes toward handling rare words. The standard approach, principal component analysis, tells a totally different story: just two directions capture 90 percent of GPT-2's activation variance, yet those same directions carry almost none of its actual function. Dimensionality also shifts depending on what's being measured. A model's confidence about its own next guess can be read from six coordinates, while its full predicted probability distribution takes hundreds.\n\nThis matters for anyone compressing, interpreting, or editing language models, because it suggests that hunting for \"important\" directions by variance alone can miss the parts that actually drive behavior. That's a quiet problem for a chunk of interpretability work that treats high-variance directions as meaningful by default. The task-weighted charts also beat both variance-based and optimal linear compression when only a few dimensions can be kept, which is exactly the situation distillation and probing tools care about.\n\nA model's internals, in other words, don't organize themselves for our convenience. They organize around whatever keeps the rare cases from breaking.","[\"interpretability\",\"language models\",\"model compression\"]","2026-09-17T04:00:00.000Z","2026-09-17T23:29:38.272Z","2026-09-17T23:29:50.194Z","published",null,[],"ai",[26,27,28],"interpretability","language models","model compression",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.18989",0,{"sections":35},[36,40,44,49,54,58,62,67,72,76,81,86,91,96],{"name":37,"slug":24,"count":38,"latest_published_at":39},"AI",3853,"2026-09-17T08:27:09.000Z",{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",648,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":18},"Hardware","hardware",154,{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",114,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":18},"Dev Tools","dev-tools",73,{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]