AI/ kv-cache compression · fine-tuning · llm research · benchmarking

Fair Learning Rates Erase a KV-Cache Compression Myth

A new study shows that giving each fine-tuning method its own learning rate wipes out the supposed edge of a popular low-rank KV-cache compression trick.

A popular trick for shrinking AI memory costs turns out to owe its reputation to a measurement error.

Researchers studying KV-cache compression, a technique that shrinks the memory language models use to track context, found a flaw in how fine-tuning recipes get compared. The method factorizes a model's key/value weights into a down-projection encoder and an up-projection decoder, then uses a short recovery fine-tune called healing to repair accuracy lost from compression. Under one shared learning rate across all three healing variants, healing only the encoder looked like a clear winner. But once each variant got its own properly tuned learning rate, that advantage vanished, and encoder-only healing matched the others, verified with three seeds, a vision-language model (Qwen2.5-VL-3B-Instruct), and two additional text-only backbones.

Encoder-only healing's real advantage isn't accuracy, it's cost: 3x fewer trainable parameters and 3x less optimizer-state memory for comparable results. That's a genuine win for anyone deploying memory-constrained models, but it's a narrower claim than the "best accuracy" result a shared-learning-rate comparison seemed to show. The warning extends well beyond KV-cache compression: any comparison of fine-tuning recipes with different parameter counts, run under one shared learning rate, risks crowning a winner that isn't real.

Call it a reminder that in machine-learning benchmarks, the method that wins is sometimes just the one that got better hyperparameters, not the better idea.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →