AI research agents - the ones that plan, search, and summarize on their own - still make things up, and a new benchmark shows exactly where.
Researchers behind a paper posted to arXiv built a framework called the PING Taxonomy, which splits hallucinations in these agents into four types: errors that propagate from earlier mistakes, errors from misreading the task's intent, errors induced by noisy search results, and errors where the agent's claims aren't grounded in any real source. They broke each agent's work into atomic actions, claims, and sub-queries, then checked every piece against fact-checking benchmarks and human review. From that process they curated DeepHalluBench, a set of 100 tasks designed to be hallucination-prone, including deliberately adversarial ones. Running six widely used deep research agents through it, every single one showed meaningful reliability gaps.
The useful part isn't the headline failure rate - it's the diagnosis. The researchers trace most of the damage to hallucination propagation, where one bad early step contaminates everything downstream, plus cognitive biases baked into how these agents plan and search. That's a more actionable finding than another leaderboard number, because it points at where in the pipeline builders should actually intervene.
Most agent benchmarks still grade the final answer and ignore the reasoning that got there, which is a bit like grading a math test only on the last line.