A new preprint claims a way to strip vision-language models of 94% of their visual tokens while barely denting accuracy.
The method, called TReVS, is outlined in a preprint titled "TReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning" (arXiv:2609.37581), posted September 30 and not yet peer-reviewed. Vision-language models burn most of their compute processing the long token strings that represent an image, and most pruning methods trim those tokens in two blind stages, first by visual redundancy, then by relevance to the text prompt. The paper's authors found that pruning before consulting the query throws away visual evidence the model needs later, so TReVS folds textual relevance into that first pass and uses a subset of attention heads the authors call "high-variance," ones more sensitive to the query, to guide a second pruning round inside the LLM's early layers. On LLaVA-1.5-7B, TReVS cut 94.4% of visual tokens while keeping 92.8% of the unpruned model's performance, which the authors say beats prior pruning methods on the same benchmark.
Token pruning is the quiet lever behind cheaper multimodal inference: every token a VLM skips is latency and cost saved at scale, which matters more as these models get embedded in chatbots, agents, and phone apps that need fast responses. What's notable is that the fix is training-free, meaning it works on an existing model like LLaVA-1.5-7B without retraining, unlike some accuracy-recovery techniques that require fine-tuning after pruning.
Still, this is one paper's numbers on one base model, not yet peer-reviewed, and "state of the art" in pruning research has a habit of getting leapfrogged within months. Take the 92.8% retained-performance figure as a snapshot, not a verdict.