AI/ ai · search-agents · llm-training · benchmarks

PrimeSeeker Trains AI Search Agents to Waste Fewer Searches

A new training method targets wasted searching instead of just adding more hops, and benchmarks show it works with far fewer tool calls.

Researchers have found a way to train AI search agents that cuts wasted searching, not just adds harder practice questions.

Most synthetic training questions for deep-search agents are scaled up crudely - add more hops, stretch the evidence trail, call it harder. A new paper argues that's an indirect proxy for the real skill: resolving a specific unnamed fact from a vague description and feeding it into the next search. The team calls this "latent anchor reasoning" and builds a framework named PrimeSeeker around it, generating questions from web-grounded anchor structures paired with a reference evidence skeleton. That skeleton guides expert training examples, then gets reused later to grade how well a reinforcement-learning policy covers the right evidence steps, using 9,221 expert trajectories to train a 30 billion parameter search agent.

Across five deep-search benchmarks, the resulting agent performed strongly, and reinforcement learning on top of the supervised model pushed it further - with low retrieval redundancy, meaning not much wasted searching. Under a fixed tool-call budget, it matched strong solution coverage while using substantially fewer tool calls than longer-horizon systems, which matters directly for anyone paying per API call to run one of these agents at scale.

Fewer tool calls per query is the kind of efficiency gain that shows up on an infrastructure bill, not just a leaderboard, and it points to where deep-search training is heading: away from brute-force longer trajectories and toward teaching the specific retrieval skill an agent actually needs. The real test is whether that discipline survives contact with messy, real-world queries beyond the five benchmarks here - if it does, capability-oriented supervision like this could become the default way search agents get trained, the same way instruction-tuning became the standard last step for making raw language models useful as chatbots.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →