A new benchmark says AI still cannot find the one paper that would have unlocked your research - and letting a model search for itself makes things worse, not better.
Researchers built ScholarCatalyst by asking 184 lead authors of 207 recent computer science papers to point to the prior work that actually advanced their own projects, each pick accompanied by a written rationale. The resulting retrieval task starts from a research question and asks a system to find those influential papers, searching only through literature that existed before the project began. Plain embedding retrieval, the kind of semantic search already common in research tools, scored 0.48 Recall@20. An agentic search setup that called the same retriever repeatedly and reasoned about what to search for scored only 0.42 - worse, not better.
That gap undercuts the assumption that wrapping an LLM agent around a search tool automatically improves results; here, the extra reasoning steps hurt retrieval. Even an agent built on Claude Fable 5.1, a model that may have already seen the completed papers during training, topped out at 0.51 Recall@20 - barely ahead of plain embeddings and nowhere near reliable.
Literature search tools have chased better recommendations for years, from citation graphs to semantic search plugins; ScholarCatalyst suggests strapping a reasoning agent onto the same retriever does not close the gap, and here it widened it.