AI agents that browse the web can be fooled by one convincing fake article - even when the truth is sitting right next to it in the search results.
Researchers built the Synthetic Web Benchmark, a fabricated mini-internet of thousands of hyperlinked pages with known-correct labels for credibility and factuality. Into that controlled environment they slipped a single high-plausibility piece of misinformation and placed it at a specific rank in the search results. Six frontier language models were then set loose to research questions inside that environment. Accuracy on those questions collapsed, the researchers report, even though the models had unlimited access to accurate sources sitting right alongside the fake one.
The bigger problem isn't that the models got fooled once - it's how they got fooled. The paper says the agents barely escalated their searches to double-check the suspicious article and ended up badly miscalibrated about how confident they should be. That's a rough combination for any agent meant to research medical, financial, or legal questions without a human checking its work.
None of this is shocking if you've watched a chatbot confidently repeat a forum post as fact, but it's the first time the failure has been isolated and measured with a controlled testbed rather than just noticed anecdotally - exactly the kind of receipt anyone pushing for AI safety testing will want on file.