AI/ ai evaluation · benchmarking · ai research

Study Finds AI Grading Rubrics Often Score the Wrong Things

A new taxonomy of nine rubric failure modes shows about one in five expert-written AI evaluation rubrics score the least important criteria the most.

A new study says the rubrics used to grade AI agents are themselves ungraded, and that's a problem.

Researchers built a taxonomy of nine ways evaluation rubrics can fail, grouped under reliability and content validity, calling it RIFT. They tested it by injecting 720 known corruptions into clean rubrics at set severity levels, then used a linear probe over measurable signals to guess which failure mode had been injected. The probe hit 75.0% accuracy, beating a frontier model asked to diagnose the same failures directly, which managed only 56.7%. Applying the method to rubrics from the GDPval and Terminal-Bench benchmarks, the team found 10 of 48 expert-written rubrics weighted their criteria backwards.

Backwards weighting is the quiet failure mode here: a rubric that scores low-stakes details at 40% and the actual point of the task at 10%. An agent can nail the one thing that matters most and still get marked down, which means some benchmark results may already be rewarding the wrong behavior without anyone noticing.

If roughly a fifth of hand-written grading criteria are miscalibrated even when experts wrote them, the automated rubric-generation tools now creeping into AI evaluation pipelines deserve at least as much skepticism, not less.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →