Compressing large language models usually means shrinking their numbers after training, and a new paper argues the math behind that process has a quiet flaw.
Post-training quantization typically breaks a model into chunks, like transformer blocks, and tunes each one separately to minimize reconstruction error, usually measured with mean squared error (MSE). Researchers tested this approach across multiple LLM families, model sizes, and quantization setups. They found that the scale of that error varies wildly from one quantization stage to the next. Because MSE translates error size directly into gradient size, early or high-error stages get much stronger optimization pushes than others, leaving the process lopsided.
That imbalance matters because quantization is the main trick making it affordable to run LLMs on everyday hardware, from phones to budget GPUs. If some layers are under-optimized simply because of how the loss function scales, the resulting compressed model performs worse than it should, for reasons that have nothing to do with model size or quantization bit-depth. The fix the researchers propose, swapping MSE for root mean squared error variants, is a drop-in change that doesn't require new hardware or retraining from scratch.
It is a reminder that in AI efficiency work, the unglamorous choice of loss function can matter as much as the headline compression ratio.