AI/ multimodal ai · model pruning · inference efficiency · ai research

Researchers Prune AI Models Path by Path to Cut Inference Cost

MWOP prunes redundant attention and FFN computation inside multimodal AI models by modality, speeding up inference with minimal accuracy loss.

Researchers have a new way to shrink multimodal AI models without gutting their accuracy: prune the computation path by path, not layer by layer.

A team proposes a technique called MWOP, short for Modality-aware Width-wise Operation Pruning, aimed at multimodal large language models - the kind that process both images and text, like LLaVA or Qwen2.5-VL. Earlier compression methods treated each attention head and each shared computation channel as one block to cut. MWOP instead found that within a single attention head, the three ways information flows - image-to-image, text-to-image, and text-to-text - carry different amounts of waste, so it prunes each path on its own. It does the same for the model's feed-forward layers, trimming visual and textual channels separately, then uses a lightweight retraining step to patch the damage, plus custom accelerator code so the finer cuts translate into actual speed, not just theory.

That distinction matters because multimodal models burn most of their compute processing long image-and-text inputs, and that cost hits cloud bills and response times directly. MWOP cuts computation per token rather than the number of tokens, so it stacks with existing token-trimming methods instead of competing with them - on LLaVA-OneVision-7B, combining the two pushed one method's prefill speedup from 2.0x to 2.9x while the model kept 99.7% of its performance across 12 benchmarks.

Those numbers come from the paper's own benchmarks on two specific models, so they're a best-case estimate until someone outside the lab reproduces them on different hardware and workloads.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →