A popular trick for shrinking AI memory costs turns out to owe its reputation to a measurement error.
Researchers studying KV-cache compression, a technique that shrinks the memory language models use to track context, found a flaw in how fine-tuning recipes get compared. The method factorizes a model's key/value weights into a down-projection encoder and an up-projection decoder, then uses a short recovery fine-tune called healing to repair accuracy lost from compression. Under one shared learning rate across all three healing variants, healing only the encoder looked like a clear winner. But once each variant got its own properly tuned learning rate, that advantage vanished, and encoder-only healing matched the others, verified with three seeds, a vision-language model (Qwen2.5-VL-3B-Instruct), and two additional text-only backbones.
Encoder-only healing's real advantage isn't accuracy, it's cost: 3x fewer trainable parameters and 3x less optimizer-state memory for comparable results. That's a genuine win for anyone deploying memory-constrained models, but it's a narrower claim than the "best accuracy" result a shared-learning-rate comparison seemed to show. The warning extends well beyond KV-cache compression: any comparison of fine-tuning recipes with different parameter counts, run under one shared learning rate, risks crowning a winner that isn't real.
Call it a reminder that in machine-learning benchmarks, the method that wins is sometimes just the one that got better hyperparameters, not the better idea.