AI/ ai · search-agents · llm-benchmarks · test-time-scaling

New Method Beats Majority Voting For AI Search Agents

TRACE ranks the search evidence behind AI agents' answers instead of just voting, beating majority voting while using far less compute.

A new selector lets AI search agents pick better final answers from a batch of parallel search attempts than simple vote-counting ever could.

Researchers introduce TRACE, a lightweight trained model that ranks completed search trajectories using the search evidence behind each candidate answer, rather than generating a fresh summary or just tallying votes. It preserves each trajectory's original evidence while sharing information across rollouts that cite the same documents. Tested across six WebQA policies and six long-horizon dataset-backbone combinations, TRACE outperformed majority voting throughout. On Qwen2.5-14B WebQA pools it scored 45.2% and 49.2% exact-match accuracy on Base and SFT models, edging out the strongest Qwen3-32B generative aggregators at 43.9% and 48.0%. On the harder FRAMES, GAIA, and BrowseComp benchmarks, it averaged 78.6% accuracy, 3.1 points above majority voting.

The bigger story is efficiency. TRACE processes queries at least 10 times faster than rival aggregators SolAgg, SummAgg, and AggAgent across all seven WebQA benchmarks. On the Base WebQA pools specifically, just 8 rollouts got within 0.4 points of what majority voting needed 64 rollouts to achieve - a real cut in the expensive search calls that make these systems costly to run.

Test-time scaling has mostly meant run more samples and vote. TRACE is a bet that smarter aggregation beats brute force - though it's one unpublished paper, and nobody outside this benchmark set has kicked the tires yet.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →