[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-hashing-fix-speeds-llm-decoding-up-to-33x":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},9514,"new-hashing-fix-speeds-llm-decoding-up-to-33x","New Hashing Fix Speeds LLM Decoding Up to 3.3x","A new technique called HHR fixes a mismatch in hash-based attention that was causing bad key retrieval, speeding up long-context LLM decoding by up to 3.3x.","A new algorithm called HHR patches a long-standing flaw in hash-based attention, and the paper backs it up with a concrete number: up to 3.3x faster decoding at 128K context length.\n\nHashing queries and keys into binary codes is a known trick for speeding up long-context LLM inference, since comparing short binary codes is cheaper than computing full attention scores. The problem is that this process throws away magnitude information, so it retrieves irrelevant keys and skips genuinely important ones. HHR, short for Hierarchical Hash Retrieval, adds two learned stages to fix that: Geometry-Aware Key Routing redistributes feature magnitudes to prune low-relevance keys more reliably, and Learned Hash Projection then re-aligns the remaining hash codes with actual query-key relevance for finer selection. Tested across multiple LLMs and benchmarks including LongBench, the combined approach lifted average scores by 1.10 points and produced up to 3.30x faster decoding and 2.83x faster end-to-end generation on Llama-3.1-8B-Instruct.\n\nLong-context inference is one of the biggest cost centers in running today's LLMs, and sparse-attention shortcuts like hashing promise speed without retraining the underlying model. What's notable here is the diagnosis, not just the fix: the researchers argue hashing's accuracy problem was never the core idea, it was a fixable engineering oversight, discarding magnitude data, rather than some fundamental limit on how well hashing can approximate attention.\n\nThe code is public on GitHub, which counts for more than any benchmark table - plenty of \"faster inference\" papers never ship anything you can actually run.","[\"llm-inference\",\"sparse-attention\",\"ai-research\",\"long-context\"]","2026-10-02T04:00:00.000Z","2026-10-02T22:43:26.024Z","2026-10-02T22:43:30.184Z","published",null,[],"ai",[26,27,28,29],"llm-inference","sparse-attention","ai-research","long-context",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.01230",0,{"sections":36},[37,40,44,48,53,57,61,66,71,76,81,86,91,96],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",5887,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",835,{"name":45,"slug":46,"count":47,"latest_published_at":18},"Policy","policy",438,{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":18},"Hardware","hardware",199,{"name":58,"slug":59,"count":60,"latest_published_at":18},"Science","science",171,{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]