[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-study-maps-14767-ai-benchmarks-and-how-they-judge-models":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},6751,"study-maps-14767-ai-benchmarks-and-how-they-judge-models","Study Maps 14,767 AI Benchmarks and How They Judge Models","A survey of nearly 15,000 papers on AI evaluation finds LLM-based grading rising fast, while AI-generated test questions have not seen the same growth.","A massive new study of AI benchmarks finds the machines are increasingly grading their own homework, but not writing it.\n\nResearchers analyzed 14,767 papers introducing or updating LLM benchmarks, published between January 2022 and August 2026. They tracked how evaluation criteria have shifted, with more tests now focused on agentic tasks, multi-step interactions, and professional work rather than simple question-answering. The study also found two diverging trends in how AI participates in the benchmarks themselves. LLM-based scoring has grown steadily across both agent-style and traditional evaluations, but model-generated test materials, meaning questions or scenarios written by AI, have not seen a similar sustained rise in recent cohorts.\n\nThat split matters because it changes who is checking whose work. If models increasingly judge other models' answers while humans still write most of the test questions, the underlying tests stay independent even as the grading doesn't. The researchers frame this as an open question, whether more AI involvement in evaluation produces more independent evidence or just launders the same models' blind spots through extra layers of automation.\n\nIt is the industry's version of grading your own exam, except so far it is only the grading pen that's automated, not the questions.","[\"ai-benchmarks\",\"llm-evaluation\",\"ai-research\",\"model-evaluation\"]","2026-09-18T04:00:00.000Z","2026-09-18T15:05:05.285Z","2026-09-18T15:05:17.226Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the dek's central claim — it says benchmarks increasingly rely on AI to 'build and grade' tests, but the source and the article's own body state model-generated test materials have NOT seen a sustained increase, only LLM-based grading has grown, so the dek contradicts the reporting.","resolved","ai",[32,33,34,35],"ai-benchmarks","llm-evaluation","ai-research","model-evaluation",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.19182",0,{"sections":42},[43,46,50,55,60,64,68,73,78,83,88,93,98,103],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",3959,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",652,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":18},"Hardware","hardware",155,{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",116,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Dev Tools","dev-tools",76,"2026-09-18T01:04:54.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]