[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-technique-cuts-ai-model-memory-use-by-32x":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8537,"new-technique-cuts-ai-model-memory-use-by-32x","New Technique Cuts AI Model Memory Use by 32x","A new arXiv preprint shows a learned selector can shrink LLM KV caches by up to 32x without hurting accuracy, letting models run cheaper at long context.","A new method can shrink the memory used by long-context AI models by up to 32 times without hurting accuracy.\n\nThe approach, called KV-Kaizen and detailed in an arXiv preprint (arXiv:2609.37988, \"KV-Kaizen: Learning Context-Adaptive Cache Compression Choices\"), targets the key-value cache that large language models use to track context. That cache can grow bigger than the model's own weights as context length increases, and since decoding is memory-bound, a bloated cache slows down every token the model generates. Instead of the common fix - evicting less-relevant tokens, which risks losing something the model needs later - KV-Kaizen trains a selector that decides, layer by layer, how to compress the cache before generation begins, using three levers: sharing cache across layers, storing it at lower precision, or truncating its low-rank representation. On a 14-billion-parameter model, that got the decode-time cache down 32x with no accuracy loss, and on models 7 billion parameters and up, a 4x cache reduction cost nothing in accuracy.\n\nThat matters because cache size, not raw compute, is often what caps how fast and how cheaply a model can process long documents, chat histories, or codebases. The paper also reports that a compressed model outperforms a smaller uncompressed model using the same cache budget, an argument for training big models first and compressing afterward rather than shrinking them from the start.\n\nIt is one preprint, not a shipped product, but if the gains hold outside the paper's own benchmarks, this is a cheaper route to long-context performance than simply buying more memory.","[\"ai\",\"llm\",\"research\",\"efficiency\"]","2026-09-30T04:00:00.000Z","2026-09-30T09:50:32.922Z","2026-09-30T09:50:37.487Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Add explicit attribution — name the source as the arXiv preprint (arXiv:2609.37988, 'KV-Kaizen: Learning Context-Adaptive Cache Compression Choices') so all the compression figures are traceable to a publication rather than floating unsourced.","resolved","ai",[30,32,33,34],"llm","research","efficiency",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.37988",0,{"sections":41},[42,45,49,53,58,63,68,73,78,82,87,92,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5104,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",785,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",417,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]