[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-technique-curbs-attention-sink-errors-in-llm-quantization":10,"sections":45},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":40,"feedback":44,"feedback_at":22,"cost_usd":44,"total_tokens":44},8144,"new-technique-curbs-attention-sink-errors-in-llm-quantization","New Technique Curbs Attention Sink Errors in LLM Quantization","OASIS curbs the attention sink problem that hurts compressed language models, though its headline number deserves a skeptical look.","A new open-source technique aims to fix a subtle flaw in how AI language models handle compression, without retraining them from scratch.\n\nResearchers built OASIS, a method for stabilizing a transformer design called AttnResidual, which is more flexible at routing information between layers but tends to create attention sinks - tokens, often the very first one, that hoard disproportionate attention weight for no clear linguistic reason. Those sinks produce activation outliers that get mangled when a model's weights and activations are quantized down to lower bit precision for cheaper inference. OASIS adds explicit null routes that let the model steer that excess attention away from problem spots at both the token and layer level. The team tested it on three open-source model families - LLaMA-3.2-1B, Qwen3-0.6B, and Phi-4 - against five baseline methods across language modeling, reasoning, and long-context benchmarks, with code posted on GitHub.\n\nQuantization is how most local and edge AI deployment happens now, squeezing models onto phones and laptops by using 8-bit or 4-bit math instead of 16 or 32-bit. If attention sinks are a hidden tax on how well compressed models perform, a fix like OASIS could matter more than another leaderboard-topping model release, since it works underneath whatever model adopts it.\n\nThe paper reports OASIS cuts perplexity by 82 percent under 8-bit weight and activation quantization and lifts 4-bit reasoning accuracy by over 42 percent on average. That first number is worth raising an eyebrow at: 8-bit quantization alone is usually close to lossless, so a swing that large likely says more about how unstable the underlying AttnResidual design is to begin with than about a fix for quantization in general - a distinction worth confirming before this becomes a standard citation.","[\"ai\",\"quantization\",\"llm\",\"open-source\"]","2026-09-28T04:00:00.000Z","2026-09-28T11:18:54.453Z","2026-09-28T11:19:01.381Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the claim that all three tested backbones are 'sub-2-billion-parameter models' — Phi-4 is a ~14B-parameter model, not sub-2B, so either drop that characterization or scope it only to the two models it actually applies to (LLaMA-3.2-1B and Qwen3-0.6B).","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The claimed 82% perplexity reduction from 8-bit weight\u002F8-bit activation quantization is implausible on its face since 8-bit quantization is typically near-lossless, suggesting a factual\u002Fnumeric error that needs verification before publishing.","ai",[35,37,38,39],"quantization","llm","open-source",[41],{"name":42,"url":43},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.17887",0,{"sections":46},[47,50,54,59,64,69,73,78,83,88,93,98,102,107],{"name":48,"slug":35,"count":49,"latest_published_at":18},"AI",4798,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Security","security",762,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Deals","deals",261,"2026-09-27T15:30:35.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Hardware","hardware",188,"2026-09-27T20:46:36.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":18},"Science","science",151,{"name":74,"slug":75,"count":76,"latest_published_at":77},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":99,"slug":100,"count":96,"latest_published_at":101},"General","general","2026-09-26T17:02:42.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]