AI/ ai · benchmarks · evaluation · research

A New Survey Tries to Make Sense of AI Research Benchmarks

A new survey finds that evaluations of AI research tools, from literature reviews to peer review, rarely use comparable methods or standards.

A new survey says we don't actually have a good way to measure whether AI research assistants are any good at research.

A team of researchers reviewed the sprawling landscape of benchmarks and studies used to judge automated research systems: tools that summarize literature, generate research ideas, run experiments, draft papers, and even review other papers. The survey sorts the field into six categories: literature synthesis, research ideation, executable workflows, scholarly writing, automatic peer review, and end-to-end research. For each category, the authors compare how tasks are built, where the correct answer comes from, who or what judges the results, and how those judgments turn into a score. Their conclusion: most evaluations check very different things, even when they claim to measure the same skill.

That inconsistency matters because papers claiming an AI system beats humans or rival models at research tasks often lean on benchmarks that were never designed to be compared against each other. The survey notes that evaluator calibration is specific to what's being judged, so a system scored on correctness and one scored on originality need different yardsticks, and that something as basic as how many attempts a model gets changes its final score. Readers are rarely told which of these variables was actually in play.

In a field racing to prove AI can do science, the real finding here might be that nobody agrees yet on what doing science well looks like on a scorecard.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →