AI/ vision-language models · clip · embedding geometry · ai research

Researchers Explain Why Fixing CLIP's Modality Gap Cuts Both Ways

A geometric analysis of CLIP and SigLIP shows one embedding quirk helps classification but hurts retrieval, depending on removal method.

A new study explains why tweaking a single geometric quirk in CLIP-style models can boost one task and quietly wreck another.

Researchers examined the modality gap, the well-documented separation between where image embeddings and text embeddings sit in the shared space that contrastive vision-language models like CLIP and SigLIP learn. Across both model families, they found that a single direction accounts for 94.4 to 99.9 percent of that separation, meaning the gap is essentially one-dimensional rather than some messy multi-part effect. Breaking down the similarity score used in classification and retrieval, they show that direction plays three distinct roles depending on the task. Subtracting it acts like a class bias in zero-shot classification, distorts rankings in cross-modal retrieval, and in mixed-modal retrieval it mostly just separates items by modality, so removing it there can actually improve results.

That last point matters because engineers have spent years treating gap-closing as a blanket fix without a clear theory of when it helps or hurts. This paper supplies the missing mechanism, a geometry-derived exponent that predicts, with a Spearman correlation of 0.93 against grid search, how much correction a retrieval system actually needs. It will not replace tuning since the gains reportedly do not transfer evenly across settings, but it replaces guesswork with a formula.

The modality gap was never a bug to squash uniformly - it looks more like a dial that needs a different setting for every job, and this paper is the first real attempt at a manual for turning it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →