[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-shrink-llm-agent-memory-costs-without-losing-accuracy":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},8070,"researchers-shrink-llm-agent-memory-costs-without-losing-accuracy","Researchers Shrink LLM Agent Memory Costs Without Losing Accuracy","A new technique called ActKV trims the memory AI agents burn on long tasks by three-quarters while keeping accuracy nearly intact.","AI agents that loop through observe-think-act cycles are quietly becoming memory hogs, and a new paper proposes a fix aimed specifically at that problem.\n\nResearchers built a system called ActKV that compresses the KV cache - the running memory an LLM keeps of everything it has seen so far - by focusing on what actually drives an agent's next action, rather than treating all cached tokens equally. It identifies and retains cache entries tied to stable, recurring action patterns, uses the model's own confidence signals to adjust how much memory budget it spends as a task evolves, and organizes the whole process into standardized, hardware-friendly memory pages. On long multi-step tasks, the team reports ActKV keeps 98.53% of the accuracy of an uncompressed cache while using only 25.98% of its peak memory. It also claims 3.97 times the token throughput and 3.58 times the task throughput of the uncompressed baseline.\n\nThat distinction - action-critical memory versus everything else - matters because agent workloads, not simple chat, are what's straining inference budgets today. Every tool call and observation an agent processes gets tacked onto its context, and serving providers pay for that memory regardless of whether it helps the next decision. A method that cuts memory to roughly a quarter while nearly matching full accuracy is a real lever on serving costs, not just a benchmark exercise.\n\nThe numbers come from the paper's own long-trace benchmarks rather than a live deployment, so the real test is whether these gains hold up once agents are juggling messier, real-world tool calls instead of curated evaluation traces.","[\"llm agents\",\"kv cache\",\"inference optimization\",\"ai research\"]","2026-09-28T04:00:00.000Z","2026-09-28T08:12:20.750Z","2026-09-28T08:12:26.784Z","published",null,[],"ai",[26,27,28,29],"llm agents","kv cache","inference optimization","ai research",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.31395",0,{"sections":36},[37,40,44,49,54,59,63,68,73,78,83,88,92,97],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",4750,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",759,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",261,"2026-09-27T15:30:35.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Hardware","hardware",188,"2026-09-27T20:46:36.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":18},"Science","science",151,{"name":64,"slug":65,"count":66,"latest_published_at":67},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":89,"slug":90,"count":86,"latest_published_at":91},"General","general","2026-09-26T17:02:42.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]