AI/ ai agents · llm evaluation · ai safety

New Test Predicts When AI Agents Fall for One-Sided Evidence

Researchers built a quick five-document test that reliably predicts whether an AI agent will be swayed by one-sided evidence in a full 45-document context.

A new audit can predict, in minutes, whether an AI agent will get swayed by lopsided evidence before you ever see its final answer.

Researchers built the test by showing agents two mirrored five-document sets and measuring how much six downstream decisions shifted between them. They froze the protocol before running it on three held-out open-weight model families, covering 18 model-task combinations. That cheap five-document signal predicted the agent's behavior in a separate, much larger 45-document test with a Spearman correlation of .855, cut average prediction error by 62% compared to assuming no effect, and got the direction right in 12 of 13 cases that actually mattered. Control experiments confirmed the cause: picking one-sided but individually ordinary documents, not just reordering the same ones, is what tips a susceptible model.

That matters because evidence selection, not just model training, is where bias can sneak into agent decisions. In a separate transfer test spanning seven open-weight model families, the same susceptibility carried over from an interactive feed-style setup to a static research dossier, and simply warning the agent that its sources might be skewed did not reliably fix it. Three deployed Codex agent tiers also ranked consistently on the cheap audit, though no single full-context effect held up after statistical correction.

All of this happened in one synthetic remote-work scenario, so treat it as a useful triage tool for testing agents, not proof that every model everywhere folds under a stacked deck.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →