AI agents are getting good at pretend research. Can they do the real thing?
HyperBrowseComp is a new benchmark built to find out. It contains 423 questions across 13 languages, each written and validated by native or highly proficient speakers. The questions are engineered to resist easy answers: finding the right one means chasing multi-step clue chains through videos, scanned documents, images, and maps rather than pulling a fact from memory. To keep things honest, the researchers first ran candidate questions past models with no internet access and tossed out any that could be answered from training data alone. They then tested several models using both provider-native search tools and a shared external retrieval system, under one common evaluation protocol, and ran a parallel human evaluation on a sample of questions.
That filtering step is the interesting part. Most browsing benchmarks eventually get memorized or gamed as models absorb the answers during training; building in a parametric-knowledge filter from the start is a direct response to that failure mode. Testing across 13 languages also pushes past the English-only assumption baked into a lot of agent benchmarks, which matters since real web research increasingly means reading sources no translation layer has touched yet.
The abstract doesn't say how any model actually performed, human or otherwise. For a benchmark explicitly designed to be "extremely challenging," that's the number worth waiting for.