[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-leapquant-shrinks-llm-memory-state-speeds-inference-147x":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8666,"leapquant-shrinks-llm-memory-state-speeds-inference-147x","LeapQuant Shrinks LLM Memory State, Speeds Inference 1.47x","A new 8-bit quantization method speeds up linear-attention kernels by up to 3.7x, though real-world inference only gets 1.47x faster.","A new quantization method compresses the memory that powers fast, linear-attention AI models down to 8 bits without losing accuracy.\n\nThe method, called LeapQuant, targets a bottleneck in linear-attention architectures like Gated DeltaNet and Kimi Delta Attention, which hybrid LLMs increasingly use to handle long context cheaply by compressing conversation history into a fixed-size recurrent state. Constantly reading and updating that state at full precision is expensive, and naive quantization degrades quality because errors pile up and a few outlier values dominate. LeapQuant quantizes the state only once per window of tokens instead of after every update, and keeps the state's largest outliers in high precision so rounding errors don't compound. Tested training-free across the Qwen, Kimi, and GLM model families, it matched FP32 baseline accuracy while cutting memory and compute costs.\n\nThe kernel-level math sped up 2.05x to 3.70x, but once you account for everything else involved in running a model, the real-world gain on Nvidia's B200, RTX PRO 6000, and RTX 5090 GPUs settled at 1.47x end-to-end. That gap between headline kernel numbers and actual inference speed is exactly the kind of detail that gets flattened in press coverage, and it's a useful reminder that a faster core operation doesn't automatically mean a faster model.\n\nEfficient long-context inference is turning into as big a battleground as raw model quality, and shaving memory costs without retraining is the unglamorous work that actually ships.","[\"linear-attention\",\"quantization\",\"llm-inference\",\"gpu\"]","2026-09-30T04:00:00.000Z","2026-09-30T18:09:55.869Z","2026-09-30T18:10:01.872Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek states 'up to 3.7x faster inference' but the body clarifies that 3.7x is only the kernel-level speedup while actual end-to-end inference speedup is 1.47x — rewrite the dek\u002Fheadline to distinguish kernel-level from end-to-end figures so it isn't contradicted by the body.","resolved","ai",[32,33,34,35],"linear-attention","quantization","llm-inference","gpu",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38166",0,{"sections":42},[43,46,50,54,59,64,68,73,78,82,87,92,97,102],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5180,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",791,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",417,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",155,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]