[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-models-can-now-guess-what-they-will-need-to-remember":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},10649,"ai-models-can-now-guess-what-they-will-need-to-remember","AI Models Can Now Guess What They Will Need to Remember","A training-free method has AI models draft short guesses at their own future answers to decide which cached prompt tokens are worth keeping.","Researchers have found a cheap way for AI models to guess what they'll need to remember later, and a single guess captures most of the benefit.\n\nA new paper introduces LORE-KV, a method for trimming an AI model's key-value cache, the working memory it uses to avoid reprocessing earlier text. Instead of judging importance based only on what the model already read, LORE-KV has the frozen model generate a few short, disposable continuations of its own likely answer, then uses those drafts to score which prompt tokens the real answer will actually draw on. The draft continuations are thrown away before the model produces its final response, so no extra model or retraining is required. On Qwen2.5-14B, the method lifted the LongBench benchmark average from 45.49 to 48.24 at a 128-token cache budget; on Mistral-7B, it raised the 16K RULER benchmark average from 45.20 to 51.05 at the same budget.\n\nLong conversations and long documents strain an AI model's memory, and most existing fixes either retrain the model or compress the past with no sense of what's coming next. LORE-KV gets an approximate view of the future for free, just by asking the model to imagine it. That's a meaningfully different bet than the synthetic-query tricks other future-aware eviction methods have leaned on.\n\nAn ablation on the 128-token cache budget found that a single imagined continuation already captures about 89% of the benefit; extra guesses barely add more. The tradeoff: generating that one guess still costs 1.5 to nearly 3 times as long as a comparable baseline's compression step, the advantage fades at bigger cache budgets, and some individual tasks get worse even as the averages improve.","[\"ai\",\"llm inference\",\"kv cache\",\"model efficiency\"]","2026-10-07T04:00:00.000Z","2026-10-09T01:21:25.233Z","2026-10-09T01:21:29.812Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The source attaches the 'single continuation recovers ~89% of the gain' ablation result generically to the B=128 setting without specifying which benchmark it applies to, but the draft pins it specifically to the Mistral-7B\u002FRULER result — rewrite to either confirm which experiment the 89% figure belongs to or state it without tying it to the Mistral-7B number.","resolved","ai",[30,32,33,34],"llm inference","kv cache","model efficiency",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.07643",0,{"sections":41},[42,46,51,56,61,66,70,75,80,84,89,94,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",6506,"2026-10-07T18:45:00.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":50},"Security","security",911,"2026-10-07T19:53:42.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",474,"2026-10-07T18:23:21.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",453,"2026-10-07T23:58:31.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",222,"2026-10-07T21:19:54.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":18},"Science","science",187,{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",174,"2026-10-07T17:41:41.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Software","software",113,"2026-10-07T18:10:00.000Z",{"name":81,"slug":82,"count":78,"latest_published_at":83},"Startups","startups","2026-10-07T23:36:57.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Dev Tools","dev-tools",105,"2026-10-07T16:59:11.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"General","general",61,"2026-10-07T22:00:24.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Gaming","gaming",56,"2026-10-07T12:00:00.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",33,"2026-10-05T11:57:17.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]