[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-damp-quantization-cuts-ai-model-memory-use-by-nearly-70":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},9329,"damp-quantization-cuts-ai-model-memory-use-by-nearly-70","DAMP Quantization Cuts AI Model Memory Use by Nearly 70%","A new quantization method called DAMP trims memory use in long-context AI models by 69% while keeping reasoning accuracy close to full precision.","A new compression trick squeezes the memory-hungry state that some AI models carry between tokens, without wrecking their reasoning.\n\nResearchers studied hybrid AI models that mix standard Softmax Attention with recurrent-state methods like Gated DeltaNet and Kimi Delta Attention, which labs use to keep long-context inference from eating too much GPU memory. Those recurrent states are normally stored in full 32-bit precision, which is heavy on memory and bandwidth. The team found that compressing them with standard INT8 or FP8 hurt complex reasoning, and more aggressive INT4 or NVFP4 compression collapsed accuracy almost entirely. Their fix, DAMP (short for Decay-Aware Mixed-Precision quantization), identifies the specific channels where compression errors matter most and keeps only those in higher-precision FP16, dropping the rest to INT8.\n\nTested on three models (Qwen3.6-35B, Kimi-Linear-48B, and Kimi-K3) across six reasoning and code benchmarks, DAMP cut recurrent-state storage by 69.1%, sped up the state-update step by as much as 2.59x, and trimmed total time per output token by up to 19%, all while holding accuracy near the uncompressed FP32 baseline. That is a real lever for anyone running long-context or agentic AI workloads, where memory bandwidth, not raw compute, is usually what slows things down.\n\nIt is a fix for a specific architecture, not a universal AI breakthrough, but if the numbers hold outside the paper, this is the kind of unglamorous engineering that quietly ends up in every serving stack.","[\"ai\",\"quantization\",\"model-efficiency\",\"inference\"]","2026-10-01T04:00:00.000Z","2026-10-02T09:48:09.935Z","2026-10-02T09:48:11.496Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"DAMP is used throughout the headline, dek, and body as a named method but its acronym is never defined — add a line spelling out 'Decay-Aware Mixed-Precision' (per the source) before or at its first use.","resolved","ai",[30,32,33,34],"quantization","model-efficiency","inference",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.27513",0,{"sections":41},[42,46,51,56,61,66,70,75,80,84,88,93,98,103],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",5694,"2026-10-01T12:05:27.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":50},"Security","security",821,"2026-10-01T14:00:00.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",431,"2026-10-01T11:08:42.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",311,"2026-10-01T14:19:00.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",197,"2026-10-01T11:37:06.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":18},"Science","science",165,{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",150,"2026-10-01T11:59:27.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":81,"slug":82,"count":78,"latest_published_at":83},"Software","software","2026-09-30T21:41:11.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":50},"Startups","startups",86,{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]