A new open-source technique aims to fix a subtle flaw in how AI language models handle compression, without retraining them from scratch.
Researchers built OASIS, a method for stabilizing a transformer design called AttnResidual, which is more flexible at routing information between layers but tends to create attention sinks - tokens, often the very first one, that hoard disproportionate attention weight for no clear linguistic reason. Those sinks produce activation outliers that get mangled when a model's weights and activations are quantized down to lower bit precision for cheaper inference. OASIS adds explicit null routes that let the model steer that excess attention away from problem spots at both the token and layer level. The team tested it on three open-source model families - LLaMA-3.2-1B, Qwen3-0.6B, and Phi-4 - against five baseline methods across language modeling, reasoning, and long-context benchmarks, with code posted on GitHub.
Quantization is how most local and edge AI deployment happens now, squeezing models onto phones and laptops by using 8-bit or 4-bit math instead of 16 or 32-bit. If attention sinks are a hidden tax on how well compressed models perform, a fix like OASIS could matter more than another leaderboard-topping model release, since it works underneath whatever model adopts it.
The paper reports OASIS cuts perplexity by 82 percent under 8-bit weight and activation quantization and lifts 4-bit reasoning accuracy by over 42 percent on average. That first number is worth raising an eyebrow at: 8-bit quantization alone is usually close to lossless, so a swing that large likely says more about how unstable the underlying AttnResidual design is to begin with than about a fix for quantization in general - a distinction worth confirming before this becomes a standard citation.