[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-method-shrinks-ai-cache-compaction-time-by-257x":10,"sections":39},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":34,"feedback":38,"feedback_at":22,"cost_usd":38,"total_tokens":38},8491,"new-method-shrinks-ai-cache-compaction-time-by-257x","New method shrinks AI cache compaction time by 25.7x","ARC-KV compresses AI model KV caches 25.7x faster than a prior method while slightly boosting accuracy, tackling a major long-context inference bottleneck.","A new technique cuts the time it takes to shrink a large language model's memory cache by 25.7x, without sacrificing accuracy.\n\nResearchers describe ARC-KV, a method for compacting the KV cache: the memory a large language model builds up so it doesn't have to re-read text it has already processed. That cache grows with every token and becomes a serious bottleneck for long, reused context, like a lengthy document a chatbot references across many separate questions. Existing reconstruction-based compaction methods, such as Attention Matching, keep accuracy high but spend most of their time on a slow, iterative search for which cache entries to keep as anchors. ARC-KV replaces that search with a trained indexer that picks anchors in a single pass, then merges and fits the remaining cache against the full version.\n\nTested on Llama-3.1-8B-Instruct across the QuALITY, RULER, and LongBench benchmarks, ARC-KV beat most reported compaction methods. At 10% cache retention on QuALITY, it nudged accuracy up slightly, from 0.6409 to 0.6474 versus Attention Matching, while cutting compaction time by 25.7x, from 959.8 seconds to 37.3 seconds. For any product that reuses a long context across many queries, such as a chatbot repeatedly consulting the same document, that's the difference between compaction running as a background cost and becoming a bottleneck.\n\nThe gains come from a single 8B model in a paper that hasn't been peer reviewed, so treat the numbers as a promising benchmark result, not a shipped product.","[\"ai\",\"llm inference\",\"kv cache compression\"]","2026-09-30T04:00:00.000Z","2026-09-30T06:59:57.148Z","2026-09-30T07:00:02.374Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Reconcile the speedup figure across the piece: the headline says 25x, the dek says '26 times faster,' and the body says 'more than 25 times' — pick one consistent, accurate figure (25.7x, per the source's 959.8s to 37.3s) and use it throughout.","resolved","ai",[30,32,33],"llm inference","kv cache compression",[35],{"name":36,"url":37},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.36835",0,{"sections":40},[41,44,48,52,57,62,67,72,77,82,87,92,97,102],{"name":42,"slug":30,"count":43,"latest_published_at":18},"AI",5028,{"name":45,"slug":46,"count":47,"latest_published_at":18},"Security","security",780,{"name":49,"slug":50,"count":51,"latest_published_at":18},"Policy","policy",417,{"name":53,"slug":54,"count":55,"latest_published_at":56},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":61},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":66},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Dev Tools","dev-tools",89,"2026-09-29T17:15:00.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]