A verification label on AI study guides mattered more than the AI itself.
Researchers ran a difference-in-differences study in a required first-year economics course at a university, splitting 170 students (340 exam marks) into two cohorts. One cohort got AI-generated podcasts, FAQs, and quiz guides built with a source-grounded model, but every item was checked by a named graduate teaching assistant before release. Students with access scored 2.34 marks higher on average on a 50-mark component, and the share of marks falling below the upper-second grade boundary dropped by 24.7 percentage points compared with the untreated cohort. The gains were statistically significant only in a narrow band, 23 to 31 marks, and roughly three-quarters of the average benefit came from the bottom-scoring fifth of students.
Average effects hide who actually benefits, and this study shows the benefit was concentrated among struggling students, not spread evenly across the class. That matters because much of the debate about AI tutoring assumes it helps everyone equally, or mainly benefits already-strong students who know how to use it well; here verification seems to replace the judgment weaker students lack. Interviews with 36 students suggest the verified label encouraged them to use the material without making them turn off their critical thinking entirely.
It's a small, single-course study, but it points to the real design question for AI in education: not generate or not, but who checks the output before a student ever sees it - and evaluations that only report averages will keep missing exactly this kind of effect.