[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-method-slashes-ai-memory-bottleneck-speeds-inference-9x":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},7950,"new-method-slashes-ai-memory-bottleneck-speeds-inference-9x","New Method Slashes AI Memory Bottleneck, Speeds Inference 9x","A new selection method picks which cached tokens matter before scoring attention, cutting retrieval overhead over 90% with under 1% accuracy loss.","A new selection method decides which cached data an AI model actually needs before it starts scoring attention, not after, and it barely dents accuracy.\n\nResearchers behind a paper called Pre-hoc Sparsity, or PrHS, tackle a well known slowdown in large language models: as conversations or documents get longer, the key-value cache that stores past context balloons, and attending over all of it gets expensive. Most existing shortcuts pick which cached tokens to keep by looking at attention scores after the fact, a method the authors say systematically misses important tokens and hurts long-range reasoning. PrHS instead selects entries before attention runs, using a mathematical bound tied to how much attention mass gets discarded, so the trade-off between speed and accuracy can be set in advance rather than discovered by trial and error. Tested on LLaMA and Mistral models, the method cut retrieval overhead by more than 90 percent on the GSM8K and CoQA benchmarks, held accuracy loss under 1 percent on the broader LongBench benchmark, and ran 9.9 times faster on the attention step while pushing 2.8 times more throughput on Nvidia A100 GPUs than standard dense attention.\n\nThe real story isn't the speedup, since plenty of sparse-attention papers claim big numbers. It's the shift from guessing after the fact to setting a hard limit in advance. That gives engineers a dial for accuracy loss instead of a hope, which matters for anyone running long-context chatbots or document analysis at scale.\n\nIt is still a paper, not a shipped inference engine, and the real test will be whether these bounds hold up outside LLaMA and Mistral and inside someone's production stack.","[\"ai\",\"llm inference\",\"kv cache\",\"attention\"]","2026-09-25T04:00:00.000Z","2026-09-26T08:16:52.986Z","2026-09-26T08:16:59.305Z","published",null,[],"ai",[24,26,27,28],"llm inference","kv cache","attention",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2602.08329",0,{"sections":35},[36,40,45,50,55,60,64,69,74,79,84,89,94,99],{"name":37,"slug":24,"count":38,"latest_published_at":39},"AI",4624,"2026-09-25T21:57:05.000Z",{"name":41,"slug":42,"count":43,"latest_published_at":44},"Security","security",748,"2026-09-26T01:30:00.000Z",{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",392,"2026-09-25T18:44:30.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",258,"2026-09-26T09:00:00.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Hardware","hardware",185,"2026-09-25T15:00:22.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":54},"Science","science",144,{"name":65,"slug":66,"count":67,"latest_published_at":68},"Consumer Tech","consumer-tech",133,"2026-09-26T07:30:06.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Software","software",90,"2026-09-25T20:55:00.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"General","general",46,"2026-09-25T02:12:57.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]