AI/ quantization · llm inference · model compression · small language models

A Simple Math Trick Makes Quantized AI Models More Accurate

A post-training tweak to how AI models predict the next word cuts quantization errors and even speeds up generation, with no retraining required.

A new paper shows you can shrink an AI model's most expensive layer without retraining it or slowing it down.

Researchers describe a post-training method called softmax reparameterization for the output head, the final layer that converts a model's internal state into probabilities across its entire vocabulary. Before quantizing that layer down to lower precision, the method subtracts a scalar multiple of the vocabulary's average row from every output row, choosing that scalar through a quick search against validation data. The trick works across three quantization approaches (RTN, activation-weighted MSE, and full-Hessian GPTQ) and includes a rank-one fix for models that use nonlinear tricks like soft-capping. On Phi-4-mini, it cut a key error metric, KL divergence between the compressed and original outputs, from 0.936 down to 0.256 at 4-bit precision, and coefficients tuned on one text dataset transferred cleanly to others.

Output heads eat a disproportionate share of inference cost in small language models because vocabularies keep growing, so any free win there compounds across every model that ships on a phone or a laptop. This method adds no extra computation at inference time for compatible heads, and it delivered a 10.8% latency cut on the Phi head with the rest of the model still running at full precision.

It's a one-dimensional search, not a new architecture, which is exactly why it's easy to imagine labs bolting it onto existing quantization pipelines instead of waiting for the next flashy compression breakthrough.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →