AI/ ai agents · llm benchmarks · content moderation · ai hallucination

AI Web Agents Skip the Investigation, Study Finds

A new benchmark finds AI agents built for moderation and policy tasks rarely dig up the hidden evidence that should change their verdict.

AI agents built to moderate content or enforce policy are supposed to dig for facts before ruling. A new benchmark suggests most don't bother.

Researchers built MIRAGE, a 750-task benchmark spanning Wikipedia edit disputes, shopping-admin adjudication, and Reddit moderation. Each task gives an agent a visible surface story that often points toward the wrong call, plus hidden evidence buried elsewhere that only shows up if the agent goes looking for it. Testing eight LLM agents across two model generations, the researchers found the agents usually navigate to the right page but fail to pull out the evidence that would flip their decision. Procedural hints, explicit nudges to go investigate, improved how much digging agents did, but on Wikipedia tasks that didn't reliably improve final verdicts, especially when the buried evidence contradicted the surface impression. Worse, 12.6% of trajectories cited facts the agents made up entirely.

That gap matters because these are exactly the agents companies want handling moderation queues and admin adjudication at scale, where the whole point is catching cases that aren't what they look like on the surface. An agent that reaches the right page but reads it wrong, or invents a supporting fact, doesn't just get the call wrong, it does so with the same confident tone as a correct one. The researchers found these failure patterns held steady across model size, generation, and reasoning architecture, which suggests better prompting alone won't fix it.

Call it the difference between finding a document and reading it. Until agents reliably do both, "investigate before deciding" is still mostly aspirational.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →