A new study explains why tweaking a single geometric quirk in CLIP-style models can boost one task and quietly wreck another.
Researchers examined the modality gap, the well-documented separation between where image embeddings and text embeddings sit in the shared space that contrastive vision-language models like CLIP and SigLIP learn. Across both model families, they found that a single direction accounts for 94.4 to 99.9 percent of that separation, meaning the gap is essentially one-dimensional rather than some messy multi-part effect. Breaking down the similarity score used in classification and retrieval, they show that direction plays three distinct roles depending on the task. Subtracting it acts like a class bias in zero-shot classification, distorts rankings in cross-modal retrieval, and in mixed-modal retrieval it mostly just separates items by modality, so removing it there can actually improve results.
That last point matters because engineers have spent years treating gap-closing as a blanket fix without a clear theory of when it helps or hurts. This paper supplies the missing mechanism, a geometry-derived exponent that predicts, with a Spearman correlation of 0.93 against grid search, how much correction a retrieval system actually needs. It will not replace tuning since the gains reportedly do not transfer evenly across settings, but it replaces guesswork with a formula.
The modality gap was never a bug to squash uniformly - it looks more like a dial that needs a different setting for every job, and this paper is the first real attempt at a manual for turning it.