Researchers built a quick pre-deployment check that predicts whether trimming tokens from a compressed CLIP model will help or wreck its accuracy on the images it already struggles with.
The team tested semantic masking, a common trick for shrinking CLIP by pruning irrelevant image tokens, across eight benchmarks built to expose spurious correlations. The effect on worst-group accuracy swung wildly: masking boosted it by as much as 82.5 percent in relative terms on some datasets and cut it by up to 100 percent on others. The culprit, they found, is what they call spurious inversion: when a dataset's misleading attribute sits in the background, CLIP's text-similarity score sometimes rates that background higher than the actual object, flipping the assumption every text- and attention-guided pruning method relies on. They introduce the Spurious Inversion Metric, a label-free check run before deployment, which predicted the direction of masking's effect with statistical significance across all eight datasets and held up across six different CLIP architectures.
That matters because naive masking was, on its own, the worst-performing method the researchers tested and also the slowest, with per-image segmentation adding up to 3.5 times baseline runtime; their new batched GPU segmentation routine cuts that overhead to 1.75 times. Gating masking by the metric's sign, instead of applying it blindly, recovered its benefits on seven of eight datasets while avoiding its worst failures. That is a real win for anyone compressing CLIP for edge or mobile deployment, where a speed-driven pruning choice can quietly tank accuracy on exactly the subgroups nobody tested.
Compression work gets benchmarked on average accuracy constantly; this is one of the few checks for whether it is quietly failing the groups that matter most.