A new calibration technique makes it dramatically cheaper to shrink big language models down to 4-bit precision.
ThinQuant is a new method for rotation learning, the technique that smooths extreme values in a model's activations so it can run at very low bit widths without breaking down. Instead of needing huge amounts of calibration data like prior gradient-free methods such as DartQuant, it picks a much smaller, geometrically representative set of activations and solves a reduced optimization problem with an iterative algorithm. On Llama-3-70B, ThinQuant finished calibration in under 12 minutes and reached a WikiText-2 perplexity of 5.63, compared with 111 minutes and a perplexity of 7.55 for DartQuant. It also scaled to Llama-3.1-405B on a single H200 GPU, a size neither SpinQuant nor DartQuant reportedly manages, finishing calibration in just over two hours with a perplexity of 2.97 versus 3.48 for a rival approach called GPTAQ+QuaRoT.
Quantization is what lets a model trained on a data center cluster run on a single GPU, but the calibration step has been a bottleneck keeping that benefit out of reach for the largest models. Making a 405-billion-parameter model quantizable on one GPU, in hours rather than not at all, turns a research curiosity into something engineers can actually deploy.
Lower perplexity scores are a solid signal, not a guarantee the compressed model still reasons or codes as well as the original, so treat these numbers as promising until someone benchmarks it on real tasks.