AI research agents that cite their sources can still steer you wrong, because the evidence they read is not a random sample.
Researchers built a method called Causal Evidence Selection Correction (CESS) to audit "Deep Research" agents, the AI tools that search documents and write cited reports. The problem: citation checks confirm a cited source backs a specific claim, but they do not check whether the agent's search exposed a representative slice of all the documents available, since early results redirect later queries and when the agent stops looking. CESS predicts the evidence direction of each candidate document and corrects for that sampling bias using the logged probabilities of which documents got picked and when the search ended. Tested on the MS2 systematic-review benchmark and on trajectories from a public Open Deep Research agent, CESS cut the average error of evidence estimates by 9.2% and 60.1% respectively, compared to simply averaging the documents the agent actually read, and shrank how much the estimate swung under opposing document rankings by 39.4% and 87.2%.
As more tools lean on agentic search to produce cited summaries, a correctly cited report can still be systematically skewed if the agent happened to read a lopsided set of sources first. The researchers also flag a separate point: correcting for that sampling bias is a different problem from measuring what happens if you change the search strategy itself, which requires an actual intervention rather than a statistical correction after the fact.
Citations were supposed to be the trust signal for AI research tools. This paper is a reminder that well-cited and well-sampled are not the same thing.