A massive new study of AI benchmarks finds the machines are increasingly grading their own homework, but not writing it.
Researchers analyzed 14,767 papers introducing or updating LLM benchmarks, published between January 2022 and August 2026. They tracked how evaluation criteria have shifted, with more tests now focused on agentic tasks, multi-step interactions, and professional work rather than simple question-answering. The study also found two diverging trends in how AI participates in the benchmarks themselves. LLM-based scoring has grown steadily across both agent-style and traditional evaluations, but model-generated test materials, meaning questions or scenarios written by AI, have not seen a similar sustained rise in recent cohorts.
That split matters because it changes who is checking whose work. If models increasingly judge other models' answers while humans still write most of the test questions, the underlying tests stay independent even as the grading doesn't. The researchers frame this as an open question, whether more AI involvement in evaluation produces more independent evidence or just launders the same models' blind spots through extra layers of automation.
It is the industry's version of grading your own exam, except so far it is only the grading pen that's automated, not the questions.