AI/ ai · machine-learning · computer-vision · clip

New Image-Text Training Trick Fixes CLIP's Modality Split

ITO adds a training-only fusion step to image-text pretraining, pushing encoders toward shared meaning and beating CLIP at modest data scales.

Researchers have found a fix for a subtle flaw in how AI models learn from image-caption pairs: the resulting embeddings often group by modality - image or text - rather than by actual meaning.

The method, called ITO, pairs two mechanisms. Multi-view alignment generates several augmented versions of each image and builds richer cross-modal matches against the paired text, which the paper credits as the main source of improvement. A separate, lightweight fusion module runs only during training, nudging the image and text encoders to produce features that work well when merged, then gets discarded before deployment, so inference still runs on the same fast dual-encoder setup as CLIP. Tested across pretraining runs from millions up to a billion image-text pairs, ITO beat standard contrastive baselines and outperformed CLIP under identical data, backbones, and compute at the 100M-1B pair range, on classification, retrieval, and multimodal benchmarks.

CLIP-style dual encoders are the backbone behind a huge share of today's vision-language tools, from search to captioning to multimodal chatbots, so a training recipe that improves them without adding inference cost is a meaningful, low-risk upgrade path. The more interesting finding is that the fusion module's benefit grows as data scale increases and appears to stabilize optimization, countering the overfitting that aggressive contrastive training tends to produce late in a run - a problem most teams just tolerate rather than fix.

The code is public on GitHub, which is more than most pretraining papers offer, but this is a preprint (v4) rather than a deployed system - the real test is whether labs training at multi-billion-pair scale see the same gains.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →