AI agents can tell their own training data disagrees with the facts in front of them, and then confidently answer wrong anyway.
A new paper introduces ISE (Identify, Solve, Escalate), a framework for scoring what the authors call epistemic humility: whether an agent notices a conflict between what it already knows and what it retrieves, works through that conflict, and flags the resulting uncertainty to the user. The researchers tested four agents across two conflict scenarios, one where the model's built-in knowledge contradicted retrieved evidence, another where two retrieved sources disagreed with each other, alongside matched tasks with no conflict at all. Agents frequently caught the contradiction in their early reasoning steps. By the final answer, though, that recognition often disappeared, replaced by a confident, uncaveated, and sometimes wrong conclusion.
The unsettling finding is that accuracy and honesty are not the same metric and do not move together. Several of the most accurate agent configurations were also the worst at admitting uncertainty, which is the exact failure mode that matters most when an agent is pulling from live or conflicting sources for someone who cannot check its work. Trying to fix this by prompting or tuning models to flag uncertainty more often tended to cost accuracy, so there is no obvious free lunch here.
Most agent benchmarks still grade on whether the final answer is right, not on whether the agent, or the person reading its output, could tell it was guessing.