AI/ ai benchmarks · multi-agent systems · llm research · scientific discovery

Language Models Struggle to Guess Research Ideas From Citations

A new benchmark shows AI models rarely guess a paper's core idea from its bibliography alone, though teamwork among models lifts accuracy roughly 2.4x.

A new benchmark says most AI models can't guess what a research paper is actually about just from its citation list.

Researchers built Reconstruction, a blind test that hides the seed paper and any literature published around the same time, giving models only an anonymized, frozen bibliography and asking them to guess the paper's core idea. An independent LLM judge then checks whether the guess matches the real idea. Across six scientific domains and 643 papers, seven frontier models nailed the actual idea only about 3 to 15 percent of the time. The researchers then tried a more elaborate setup: four models cross-reviewing each other's guesses and competing in a Swiss-style tournament to pick the best hypothesis, using only the given references and no web search.

That multi-agent pipeline pushed match rates up to roughly 23 to 42 percent across all six domains - about 2.4 times better than the best single model working alone. It's a real data point in the debate over whether AI can meaningfully assist with hypothesis generation rather than just retrieval or drafting: on this evidence, one model reading a bibliography is closer to guessing than reasoning, while structured collaboration between models closes real ground.

Still, the gap between a 23-42 percent match rate and genuine scientific insight is wide, and this is an unreviewed arXiv draft, not a peer-reviewed verdict.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →