Ask an AI agent to solve a math problem, then secretly hand it a wrong answer from its own calculator, and there's a good chance it won't catch the mistake.
Researchers tested AI agents on 31 math problems using a hidden interceptor that swapped real tool outputs for plausible but incorrect ones. With no verification step in place, accuracy on corrupted problems fell from 100% to 72.4%. The team then tested three fixes: a mandatory same-context reflection step, an optional fresh-context check, and an optional structural check. Mandatory reflection fully restored accuracy to 100%, while the optional checks only helped when the model actually chose to run them.
That gap matters because agents are increasingly trusted to chain together calculators, code interpreters, and search tools without a human checking every step. The study suggests reliability isn't just about having a verification tool available - it's about whether the agent is forced to use it. Leave the choice up to the model, and you're gambling on judgment nobody tested.
A follow-up test found that restarting a corrupted problem from scratch after catching the error worked every time - a reminder that the cheapest fix for a lying tool is often just not trusting it in the first place.