[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-fair-learning-rates-erase-a-kv-cache-compression-myth":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},10935,"fair-learning-rates-erase-a-kv-cache-compression-myth","Fair Learning Rates Erase a KV-Cache Compression Myth","A new study shows that giving each fine-tuning method its own learning rate wipes out the supposed edge of a popular low-rank KV-cache compression trick.","A popular trick for shrinking AI memory costs turns out to owe its reputation to a measurement error.\n\nResearchers studying KV-cache compression, a technique that shrinks the memory language models use to track context, found a flaw in how fine-tuning recipes get compared. The method factorizes a model's key\u002Fvalue weights into a down-projection encoder and an up-projection decoder, then uses a short recovery fine-tune called healing to repair accuracy lost from compression. Under one shared learning rate across all three healing variants, healing only the encoder looked like a clear winner. But once each variant got its own properly tuned learning rate, that advantage vanished, and encoder-only healing matched the others, verified with three seeds, a vision-language model (Qwen2.5-VL-3B-Instruct), and two additional text-only backbones.\n\nEncoder-only healing's real advantage isn't accuracy, it's cost: 3x fewer trainable parameters and 3x less optimizer-state memory for comparable results. That's a genuine win for anyone deploying memory-constrained models, but it's a narrower claim than the \"best accuracy\" result a shared-learning-rate comparison seemed to show. The warning extends well beyond KV-cache compression: any comparison of fine-tuning recipes with different parameter counts, run under one shared learning rate, risks crowning a winner that isn't real.\n\nCall it a reminder that in machine-learning benchmarks, the method that wins is sometimes just the one that got better hyperparameters, not the better idea.","[\"kv-cache compression\",\"fine-tuning\",\"llm research\",\"benchmarking\"]","2026-10-09T04:00:00.000Z","2026-10-09T22:24:49.455Z","2026-10-09T22:24:52.344Z","published",null,[],"ai",[26,27,28,29],"kv-cache compression","fine-tuning","llm research","benchmarking",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.10552",0,{"sections":36},[37,40,44,49,54,58,62,67,72,77,82,87,92,97],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",6708,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",931,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",486,"2026-10-08T22:40:11.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",474,"2026-10-08T22:00:00.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":18},"Hardware","hardware",231,{"name":59,"slug":60,"count":61,"latest_published_at":18},"Science","science",192,{"name":63,"slug":64,"count":65,"latest_published_at":66},"Consumer Tech","consumer-tech",181,"2026-10-08T23:26:35.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Startups","startups",117,"2026-10-08T16:45:00.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Software","software",114,"2026-10-08T17:57:01.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Dev Tools","dev-tools",105,"2026-10-07T16:59:11.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"General","general",66,"2026-10-09T04:46:11.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",58,"2026-10-08T20:08:45.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Reviews","reviews",34,"2026-10-08T14:00:22.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]