AI/ ai · model-compression · research · llms

A Better Way to Measure How Compression Changes AI Outputs

Researchers find that total variation, not the widely used KL divergence, accurately predicts how often compressed language models flip their answers.

A new study says the metric everyone quotes for compressed AI models measures the wrong thing.

Researchers tested 802 compressed and perturbed versions of 19 open-source language models across five text corpora and nine unrelated perturbation methods. They measured how often a compressed model's top-choice word differed from the original dense model's choice, a measure they call the flip rate. A simpler statistic called total variation predicted that flip rate almost one-to-one, with a median ratio of 1.05. KL divergence, the metric compression reports actually use, connects to flip rate only loosely, through a square root and a correction factor that varies fourfold across models and corpora.

That gap is not academic: when comparing two compressors whose flip rates differ by 10 percent or more, KL divergence picks the wrong one as 'more faithful' 11 percent of the time, versus 1 percent for total variation. The pattern held up in follow-up tests on a held-out code corpus and on new models with real compression kernels. For anyone deploying a quantized or pruned model and expecting it to act like the original, the KL number in a vendor's report may be quietly underselling how much actually changed, including in speculative decoding setups like vLLM, where total variation predicted draft-acceptance rates within 1 to 2 percent.

Compression benchmarks already lean on numbers that flatter the compressor; this is one more reason to check which distance metric is actually behind a 'minimal divergence' claim.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →