A new benchmark compression method claims it can cut vision-language model testing costs by more than 97 percent without scrambling which model comes out on top.
Researchers publishing on arXiv introduced PRIMEBench, a four-stage system for shrinking the benchmarks used to evaluate vision-language models, AI systems that answer questions about images and text together. The pipeline first strips out items that can be answered without even looking at the image, plus items every model already gets right regardless. It then picks one representative benchmark per capability category, prunes what is left using a new scoring method that weighs both how much models disagree on an item and how much the image itself matters to answering it, and finally trims down the number of surviving categories. The researchers say this approach hits its best fidelity, meaning it most reliably reproduces the original model rankings, when only 5 percent of items are kept, tested on models that had no role in designing the pruning method. The suite they ultimately released cuts even harder, discarding more than 97 percent of items overall.
That is a meaningful claim, because benchmark bloat is a real cost in AI research. As vision-language models multiply and benchmarks expand to cover more capabilities, running a full test suite on every new model burns compute and time. Compression methods that keep the ranking accurate while gutting the workload have existed for text-only language models for years; doing the same for benchmarks built around images has lagged, largely because deciding which images actually matter is harder than deciding which words do.
A benchmark that survives on roughly 3 percent of its original items is either a genuinely useful shortcut or a sign that most vision-language benchmarks were padded with redundant questions to begin with. The paper does not fully settle which, though it does note that evaluation quality shifts as the pool of models being compared grows and changes, which is its own quiet warning about trusting any fixed benchmark for too long.