AI/ ai agents · visual search · retrieval · research

AI Agents Learn to Search Images Without Writing Queries

A new retrieval system lets AI agents navigate visual search in one step instead of several, cutting token use by 61 percent in early tests.

A new retrieval tool lets AI agents jump through chains of related images in one step, instead of guessing at search terms for each one.

Researchers built VHOP, a benchmark and data generation framework for testing visual search at five difficulty levels, then trained VHOP-Router, an embedding model that retrieves multi-step image chains directly in visual latent space. The training pipeline combines supervised fine-tuning, imitation learning, and reinforcement learning. In testing, VHOP-Router lifted retrieval accuracy from under 5 percent to 76.3 percent, boosted agentic task success by 52.7 percent, and cut average token usage by 61 percent, from 1,886 tokens to 728. Compared to a baseline that retrieves the top 50 results per step, it used 23 times fewer in-context images and cut the cumulative API payload by 35 times.

The real finding here is where the bottleneck actually sits. Swapping in a better LLM agent only bought a 3.7 percent improvement, while fixing the retrieval step delivered more than ten times that gain. That is a pointed argument against the current habit of throwing a bigger model at every agentic-search problem.

The usual caveat applies: these are the authors' own benchmark and numbers, not independent replication, so treat the specific percentages as a claim to watch rather than a settled result.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →