Two AI models that never met can still find common ground, a new study shows.
Researchers tested whether independently trained models - one trained on text, another on images - naturally organize their internal representations the same way, even without any matched examples to learn from. Using a technique called Wasserstein Procrustes (a way of finding a single rotation that lines up two sets of points), they aligned embeddings from separate text and image models without showing the system a single paired example. The method worked best when a simple geometric metric predicted the two models' internal structures were already similar enough to match. With even a handful of paired examples, the approach beat existing alignment techniques, and it stayed competitive once more pairs were added.
This chips away at an assumption baked into most multimodal AI: that you need huge paired datasets - captioned images, transcribed audio - to get a text model and an image model talking to each other. If independently trained models already land in roughly the same geometric neighborhood, the labeling budget for building cross-modal systems could shrink substantially. The researchers also showed the aligned embeddings could drive text-to-image generation without any paired training data, a result that usually demands significant labeled data to even attempt.
It's not a universal fix - it still depends on the two models' geometries being compatible in the first place - but it's a tidy bit of evidence for the harder, weirder claim that separately trained AI models keep arriving at the same shape of the world.