A new benchmark doesn't just grade whether an AI research agent gets the right answer - it tracks whether the agent ever actually looked at the paper that mattered.
Researchers built a method called decision checkpoints that logs every observation and tool action an AI search agent takes while hunting through scientific literature, without peeking at the answer key until after the fact. They ran it across five different search strategies on 540 answerable questions from AutoResearchBench Deep, a benchmark stocked with papers the agents were guaranteed to be able to find. A simple keyword-search approach hit 24.6% accuracy, well ahead of a raw, unstructured search at 17.8%. But the real finding was underneath the scoreboard: keyword search produced fewer wrong answers where the agent never found the target paper at all, and more wrong answers where it found the paper, left it unopened, and guessed anyway.
Final-answer accuracy is the only number most AI benchmarks report, and it conceals the difference between "never found the evidence" and "found it, ignored it, got it wrong anyway." Two of the search strategies tested here - reading documents before searching further, versus keyword search - landed on the identical 24.6% accuracy score despite triggering target inspections on different numbers of questions (199 versus 191) and racking up 27.4% more search calls in the read-first condition. A dashboard that only shows the final percentage would call those two systems equivalent. They are not.
Call it the agent-evaluation equivalent of a doctor who diagnoses correctly half the time without ever reading the chart - the score tells you nothing about whether to trust the process.