A new retrieval tool lets AI agents jump through chains of related images in one step, instead of guessing at search terms for each one.
Researchers built VHOP, a benchmark and data generation framework for testing visual search at five difficulty levels, then trained VHOP-Router, an embedding model that retrieves multi-step image chains directly in visual latent space. The training pipeline combines supervised fine-tuning, imitation learning, and reinforcement learning. In testing, VHOP-Router lifted retrieval accuracy from under 5 percent to 76.3 percent, boosted agentic task success by 52.7 percent, and cut average token usage by 61 percent, from 1,886 tokens to 728. Compared to a baseline that retrieves the top 50 results per step, it used 23 times fewer in-context images and cut the cumulative API payload by 35 times.
The real finding here is where the bottleneck actually sits. Swapping in a better LLM agent only bought a 3.7 percent improvement, while fixing the retrieval step delivered more than ten times that gain. That is a pointed argument against the current habit of throwing a bigger model at every agentic-search problem.
The usual caveat applies: these are the authors' own benchmark and numbers, not independent replication, so treat the specific percentages as a claim to watch rather than a settled result.