Newer, pricier AI models are only slightly better than their predecessors at one of research's most tedious jobs: deciding which studies belong in a systematic review.
Researchers tested eight recent large language models against a software engineering screening benchmark called SESR-Eval, plus a smaller version they built for cheaper testing. The average accuracy score across review topics rose from 0.347 to 0.365 on a measure called MCC, where 1.0 is perfect and 0 is a coin flip. That is a small bump. The models mostly agreed with each other on screening calls, and tweaking the inclusion and exclusion criteria given to them produced only modest improvements in recall.
This matters because systematic reviews are the backbone of evidence-based research, and screening hundreds or thousands of papers by hand is slow and expensive. The promise of LLMs was that newer, more capable models would keep closing the gap with human reviewers. Instead, the gains have nearly flatlined even as the models themselves have gotten far more expensive to run, and the choice of which review topic you're screening still matters more than which model you pick.
It is a useful reality check after two years of assuming every new model release would make automation problems like this one disappear on its own.