A new position paper says the AI research world has been asking the wrong question about lab-bench chatbots.
Most work on LLM-based "AI Scientist" agents treats them as autonomous researchers, judged on how well they can generate hypotheses or run experiments solo. The authors argue that framing is backwards. Drawing on literature review, empirical analysis, and real-world case studies, they say the right unit of analysis is the human-agent pair, not the agent alone. They point to documented incidents and studies showing that deploying these agents without accounting for how scientists actually work with them narrows the diversity of scientific inquiry pursued.
This matters because most AI-in-science coverage treats capability benchmarks as the whole story: can the agent design a better experiment than a postdoc. The paper's case studies suggest the more consequential variable is the working relationship - how a scientist's judgment and an agent's output shape each other over a project, not just what the agent can do unsupervised. Get that dynamic wrong and you don't just get worse science, you get narrower science, with fewer research directions explored.
The authors want a new research agenda: mathematical frameworks for modeling human-AI synergy in scientific work, not just leaderboards for agent autonomy. It's a modest ask on paper, but a pointed one - it implies most current AI Scientist benchmarks are measuring the wrong thing entirely.