AI graders can reward an answer for including information that isn't actually there - and a new training trick fixes that.
Researchers built a system called MetaRubric for rubric-based reinforcement learning, a technique where a model's response earns partial credit for hitting a checklist of requirements instead of a single right-or-wrong score. They identified a failure mode they call Vacuous Credit: a rubric judge keeps awarding points for a criterion even after the specific fact or action it's checking for has been removed from the response. That bug can flip the training signal, making a worse answer score better than a correct one. MetaRubric counters it by testing responses against counterfactual versions of each prompt, built by swapping one key fact, and only granting credit when a response actually contains evidence for the claim. Between training stages, the model's own outputs are used to revise and reweight the rubric criteria.
The payoff is a real jump in accuracy, not a marginal one. MetaRubric improved PubMedQA accuracy by 6 to 20 percentage points over a standard static-judge setup, with further gains on HealthBench-Hard and two multimodal medical benchmarks. That matters because reward hacking - models learning to game the scoring rubric rather than the actual task - is one of the biggest obstacles to using reinforcement learning on messy, open-ended problems like medical question-answering.
It's a useful reminder that in reinforcement learning, the grader is often the weakest link, not the model being trained.