AI/ ai · llm · quantization · model-compression

New Quantization Method Pushes LLMs Below One Bit Per Weight

ShamAN-Q compresses LLM weights below one bit per parameter while improving perplexity over its predecessor, NanoQuant, on Qwen3-Base models.

Researchers have found a way to shrink large language models to less than one bit per parameter without losing as much quality as earlier methods did.

The technique, called ShamAN-Q, builds on an existing sub-1-bit quantization method named NanoQuant. Where NanoQuant treats each weight's reconstruction error independently, ShamAN-Q estimates a denser curvature matrix, an approach borrowed from the Shampoo optimizer, fit to a small calibration set. It solves Sylvester equations instead of NanoQuant's ADMM updates, recalculates that curvature statistic layer by layer as quantization proceeds instead of once upfront, and spreads compression rank unevenly across layers rather than applying it uniformly.

On Qwen3-Base models, ShamAN-Q lowered WikiText-2 perplexity at roughly 1 bit per weight across all three sizes tested (0.6B, 1.7B, 4B), while zero-shot accuracy on the Eleuther LM Evaluation Harness held steady or improved. The sharper number: on the 0.6B model, ShamAN-Q's perplexity at about 0.8 bits per weight matched NanoQuant's perplexity at about 1.0 bits per weight, the same quality at roughly a fifth less storage.

Sub-1-bit quantization has mostly been a research curiosity rather than something anyone ships, since quality usually falls off a cliff below 2 bits. Shaving a few more tenths of a bit without cratering perplexity is a real improvement, but it's still a lab result on open base models, not evidence this holds up on the instruction-tuned, larger models most people actually run.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →