A tweak to how vision-language models learn now lets CLIP notice details it used to miss entirely.
Researchers revisited a method called KUEA, which fine-tunes CLIP's image encoder by matching its internal structure to DINOv2, a model built for fine-grained visual recognition. They found that weakening KUEA's core alignment signal did not noticeably hurt performance, suggesting the original approach was not doing what its designers thought. So the team built a new alignment technique based on Kernel Canonical Correlation Analysis, which looks for correlated patterns across feature subspaces instead of forcing an element-by-element match, then extended it into a three-view version called 3vKCCA that also pulls in signals from CLIP's text encoder during training. On CLIP ViT-L/14, the new method raised accuracy on the MMVP-VLM fine-grained benchmark from 17.8 to 25.9, about a 45% improvement, while keeping the model's image-text search performance intact.
CLIP and models like it underpin a lot of the image search, captioning, and multimodal AI tools built over the past few years, and their blind spot for fine visual detail, counting objects, telling left from right, spotting small differences, has been a persistent limitation. This result suggests that gap has as much to do with how researchers have been training fixes as with CLIP's architecture: the earlier KUEA approach looked rigorous but, by this paper's own test, was not targeting the right thing.
A 45% gain on one benchmark is not a cure for CLIP's blind spots, and it is still an academic result with no sign yet of production use.