AI/ ai · llm-compression · model-efficiency · research

New Framework Compresses LLM Attention Without Fine-Tuning

A new compression framework shrinks LLM attention weights without fine-tuning and beats rivals on perplexity across five modern GQA models.

A new compression method shrinks the attention layers inside large language models without retraining them, and it beats existing techniques at the toughest compression settings.

Researchers built a framework called FTC (Functional Tucker Compression) that compresses an LLM's attention weights after training is finished, with no fine-tuning or gradient-based repair required. Older compression methods squeeze each attention matrix in isolation. FTC instead accounts for how compressing one part of the attention block changes what the next part sees, and it exploits shared structure between the query, key and value weights to compress them together under a fixed storage budget, while handling the output projection separately. The team tested FTC on seven decoder-only LLMs ranging from 6 billion to 32 billion parameters. On five of those seven, the ones built on the now-common grouped-query attention (GQA) design, FTC produced the lowest WikiText-2 perplexity, a standard text-prediction benchmark, of any method compared, at every compression level tested on those models, with the advantage growing under more aggressive compression.

The no-fine-tuning part is the real pitch. Compression methods that need a gradient-based recovery pass after shrinking a model add GPU time and engineering overhead, which is exactly what teams deploying on constrained hardware are trying to avoid. FTC's gains also show up most under aggressive compression, the regime where squeezing a model onto cheaper hardware usually hurts output quality the most.

It is worth remembering this only compresses attention weights, which are one slice of a model's total footprint, so the real-world savings depend on what else gets compressed alongside it. And this is a preprint, not a peer-reviewed result; a lower perplexity score is a reasonable proxy for model quality, not a guarantee that outputs will actually read better to a person.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →