AI agents that call outside tools will often believe them even when the tools are wrong.
The findings come from an arXiv preprint (arXiv:2609.37153) that introduces ToxicBench, a benchmark feeding AI data agents both clean and deliberately poisoned tool outputs across numerical, label, schema, and retrieval errors. Across 118 tasks tested on GPT-based agents with three different tool adapters, poisoned evidence dropped task success by 26 to 39 percentage points compared to clean runs. Simple retries helped when an agent hit a single bad result, but when the same error repeated, agents kept adopting the wrong answer even after checking it. The researchers validated their automated scoring against 200 human-reviewed trajectories and found 96 percent agreement on task success.
This matters because agentic data tools are being sold on the premise that they can be trusted to fetch and act on external data unsupervised. The study shows the failure is not just missing evidence, it is agents seeing contradictory evidence and choosing to believe the tool anyway, a distinct problem from ordinary hallucination that existing retry logic does not fix once errors repeat.
Anyone deploying an AI agent to query databases or APIs might want to ask not just whether it can check its work, but whether it will believe what it finds once it does.