AI/ llm quantization · ai research · model compression · efficient ai

New method speeds LLM quantization, scales to 405B model

A new rotation-based method slashes LLM quantization calibration time and scales to a 405-billion-parameter model on a single GPU.

A new calibration technique makes it dramatically cheaper to shrink big language models down to 4-bit precision.

ThinQuant is a new method for rotation learning, the technique that smooths extreme values in a model's activations so it can run at very low bit widths without breaking down. Instead of needing huge amounts of calibration data like prior gradient-free methods such as DartQuant, it picks a much smaller, geometrically representative set of activations and solves a reduced optimization problem with an iterative algorithm. On Llama-3-70B, ThinQuant finished calibration in under 12 minutes and reached a WikiText-2 perplexity of 5.63, compared with 111 minutes and a perplexity of 7.55 for DartQuant. It also scaled to Llama-3.1-405B on a single H200 GPU, a size neither SpinQuant nor DartQuant reportedly manages, finishing calibration in just over two hours with a perplexity of 2.97 versus 3.48 for a rival approach called GPTAQ+QuaRoT.

Quantization is what lets a model trained on a data center cluster run on a single GPU, but the calibration step has been a bottleneck keeping that benefit out of reach for the largest models. Making a 405-billion-parameter model quantizable on one GPU, in hours rather than not at all, turns a research curiosity into something engineers can actually deploy.

Lower perplexity scores are a solid signal, not a guarantee the compressed model still reasons or codes as well as the original, so treat these numbers as promising until someone benchmarks it on real tasks.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →