[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-valuediff-targets-llm-cache-trimming-for-weak-sink-models":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},7343,"valuediff-targets-llm-cache-trimming-for-weak-sink-models","ValueDiff Targets LLM Cache Trimming for Weak-Sink Models","ValueDiff ranks tokens by value-vector geometry, beating older cache-eviction methods on LLM architectures that suppress attention sinks.","Researchers have found a smarter way to trim the memory large language models burn through during long conversations, tuned for the newer model designs that broke earlier compression tricks.\n\nThe technique, called ValueDiff, decides which tokens to delete from a model's key-value cache (the running memory of past tokens a model checks to generate each new word) once that cache grows too large to keep in full. Older eviction methods leaned on attention sinks, a handful of tokens the model reliably keeps referring back to, using them as a proxy for what else is safe to discard. Newer architectural tricks such as QK-normalization, gated attention, learned attention sinks, and logit softcapping weaken those sinks, so the old proxy gets noisy. The researchers found that in these models, a token's value vector (the internal numbers it contributes to the model's output) tends to scatter further from the cache average than its key vector does, and that scatter turned out to be the more reliable clue for what to keep.\n\nCache size is one of the biggest costs of running long-context models, and the industry keeps shipping the very architectures, gated attention and QK-normalization among them, that undercut sink-based compression. Tested across seven sink-suppressed models on RULER, LongBench, and MATH-500, ValueDiff held onto 88-99% of full-cache accuracy at a tight 2k-token budget and beat the strongest prior method by about nine points on LongBench at 4k tokens (92% versus 83%), with gains up to roughly 20 points on math reasoning for gated-attention models specifically.\n\nIt is one arXiv preprint, not a shipped product feature, but it is a useful reminder that memory-saving tricks tuned for one generation of transformer design can quietly stop working on the next.","[\"llms\",\"kv-cache\",\"inference-efficiency\",\"ai-research\"]","2026-09-23T04:00:00.000Z","2026-09-23T08:31:22.167Z","2026-09-23T08:31:26.761Z","published",null,[],"ai",[26,27,28,29],"llms","kv-cache","inference-efficiency","ai-research",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.23314",0,{"sections":36},[37,40,44,49,54,58,62,67,72,77,82,87,92,97],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",4297,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",710,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",369,"2026-09-23T02:13:52.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",202,"2026-09-22T23:00:04.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":18},"Hardware","hardware",169,{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",133,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",110,"2026-09-22T20:00:00.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Software","software",80,"2026-09-22T23:32:52.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Dev Tools","dev-tools",79,"2026-09-22T22:21:13.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Startups","startups",65,"2026-09-22T22:06:48.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Gaming","gaming",45,"2026-09-22T15:35:06.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",43,"2026-09-21T23:48:56.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Reviews","reviews",27,"2026-09-22T13:00:00.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]