A new search framework skips tags and metadata entirely, letting you just describe what you're looking for.
Researchers have built AiSearch, a framework that uses vision language models to search images and video with plain natural language queries, requiring no task-specific training. The system supports interactive refinement, so when results miss the mark, feedback from the user nudges the next round closer to what they actually meant. AiSearch also includes a benchmarking layer that runs multiple VLMs side by side on the same query, letting users see which model handles their particular content best. The whole approach leans on the zero-shot abilities VLMs already have, rather than retraining a model for each new collection.
Most consumer and enterprise search tools still rely on manually applied tags or embeddings trained narrowly for one catalog, which breaks down the moment content doesn't fit the mold. Treating model choice as something a user picks and compares, rather than something locked in by an engineering team months earlier, is the more interesting idea here: it turns retrieval into a visible, swappable component instead of a black box. For anyone building a product on top of image or video search, that's a more honest way to ship, since the system shows its work instead of hiding behind one hardcoded model.
Of course, this is a research paper, not a shipped tool, and interactive refinement that looks smooth in a demo often strains against a messy, million-file archive with duplicate, mislabeled, or low-quality content. The real test is what happens when someone points it at exactly that.