A new paper shows you can shrink an AI model's most expensive layer without retraining it or slowing it down.
Researchers describe a post-training method called softmax reparameterization for the output head, the final layer that converts a model's internal state into probabilities across its entire vocabulary. Before quantizing that layer down to lower precision, the method subtracts a scalar multiple of the vocabulary's average row from every output row, choosing that scalar through a quick search against validation data. The trick works across three quantization approaches (RTN, activation-weighted MSE, and full-Hessian GPTQ) and includes a rank-one fix for models that use nonlinear tricks like soft-capping. On Phi-4-mini, it cut a key error metric, KL divergence between the compressed and original outputs, from 0.936 down to 0.256 at 4-bit precision, and coefficients tuned on one text dataset transferred cleanly to others.
Output heads eat a disproportionate share of inference cost in small language models because vocabularies keep growing, so any free win there compounds across every model that ships on a phone or a laptop. This method adds no extra computation at inference time for compatible heads, and it delivered a 10.8% latency cut on the Phi head with the rest of the model still running at full precision.
It's a one-dimensional search, not a new architecture, which is exactly why it's easy to imagine labs bolting it onto existing quantization pipelines instead of waiting for the next flashy compression breakthrough.