AI/ llm-quantization · model-compression · inference-optimization

New Trick Shrinks LLM Weights Without Pricey Transforms

A new quantization method compresses LLM weights at very low bit rates while keeping decoding fast and accurate, skipping costly preprocessing.

Researchers have found a cheaper way to shrink large language models without the usual accuracy tax.

The method, called XOR-Trellis, compresses model weights using trellis-coded quantization, a technique that packs weights into very few bits without the massive lookup tables that typical vector quantization needs. The paper tackles two specific headaches: decoding those compressed weights fast enough that unpacking them does not become the slow part of running the model, and keeping accuracy high without relying on expensive Hadamard-based transformations that are normally needed to make quantization errors well-behaved. The authors introduce a simplified, hardware-friendly way to map compressed states back to real values, plus a new optimization objective that accounts for which weights matter most to model behavior, computed directly in the original weight space rather than a transformed one.

This matters because quantization is the main lever for running big models on smaller, cheaper hardware, and the two costs this paper attacks, decoding speed and preprocessing overhead, are exactly what make ultra-low-bit quantization impractical in production today. Dropping the Hadamard transform step in particular removes a computational tax that most state-of-the-art low-bit methods have treated as unavoidable.

If the results hold up outside the lab, it is a sign that the quantization field is maturing past brute-force transform tricks toward leaner, purpose-built math.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →