A team of researchers built a benchmark that catches AI-generated science papers not by flagging suspicious words, but by tracing where the reasoning falls apart.
The group assembled SciSlopBench, a set of 390 AI-written papers mostly in computer science, each matched against a human-written paper tackling the same problem. They scored each pair on six measures across structure, argument, and supporting artifacts like figures and citations, and used those scores to pick out the AI-written paper with 85.9% accuracy. That beats Binoculars, an existing AI-text detector, which managed just 68.7% on the same pairs. The slop scores also tracked real-world outcomes: papers with more of these patterns got lower ratings at the ICLR machine-learning conference and were more likely to be rejected, a pattern that held in every year from 2017 to 2025.
Token-based AI detectors have been failing for a while now, easily fooled by a quick paraphrase pass. This approach instead targets something harder to fake: whether the argument actually holds together once you strip away fluent prose, which matters more as AI-assisted drafts flood journals and conference submissions that rely on already-stretched peer reviewers.
The team's fix, SciSlopHarness, only closes 63% of the gap between AI and human writing, and it only works by checking proposed edits against the paper's actual experimental records. Prompting the model directly to remove the slop just taught it to game the metric instead. Detection, it turns out, is still well ahead of correction.