AI/ quantization · llm-inference · nvfp4 · nvidia

Researchers Tune NVFP4 Quantization to Cut LLM Accuracy Loss

OSFP4 optimizes smoothing and block scales together, keeping NVFP4-compressed models more accurate while barely slowing inference speed.

A new quantization technique squeezes more accuracy out of Nvidia's 4-bit NVFP4 format for running large language models.

Researchers built OSFP4, a scheme that jointly optimizes two things quantization engineers usually tune separately: a diagonal smoothing matrix applied to each linear layer, and the block scales used to map numbers into 4-bit buckets. The method accounts for how rounding actually happens, whether standard round-to-nearest or the successive-cancellation trick used in GPTQ. To make the joint optimization tractable, the team modeled a randomized multiplicative-dither version of the FP4 quantizer rather than the fixed deterministic one most tools use. In testing, OSFP4 beat the other quantization methods it was compared against on average accuracy, while keeping roughly 94-97% of the prefill throughput of Nvidia's stock NVFP4 implementation.

NVFP4 is Nvidia's pitch for running inference on fewer, cheaper bits without wrecking model quality, and tensor cores on recent GPUs accelerate it natively. The catch has always been that aggressive 4-bit quantization degrades accuracy unless rounding and scaling are tuned carefully, and most existing tricks handle smoothing and scaling as separate steps. OSFP4's bet is that treating them as one optimization problem recovers accuracy that separate tuning leaves on the table, for only a small throughput cost.

It is a narrow, technical win, not a new model or product, but it is the kind of unglamorous engineering that decides whether 4-bit inference becomes the default for the next GPU generation or stays a brief hack.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →