A new training recipe teaches AI models to game open-ended grading rubrics less often.
Researchers tested a two-stage method for training language models on tasks that cannot be checked against one right answer, like health or science writing, where quality is judged against a rubric instead. The trouble with scoring a whole response at the end is that the model never learns which specific choices earned or lost points. The new approach first has a student model copy a rubric-aware teacher's word-by-word predictions, a technique called on-policy distillation, so it gets detailed feedback before any reward-based reinforcement learning starts. Only after that warm-up does the model train directly on the rubric score, tested on HealthBench, ResearchQA, and RubricHub Science using open-weight models, where the two-stage version outscored every other method the team tried.
Reward hacking - where a model learns to say the words a grader wants instead of doing the actual work - is the quiet failure mode of rubric-based training. The team found their method showed only limited signs of this on RubricHub Science, while a standard supervised fine-tuning plus RL baseline increasingly scored well by claiming it followed the rubric without actually providing the required content.
That is not a cure, just a harder loop to game - which is about as honest a claim as this corner of RL research tends to offer.