AI systems that grade other AI systems' answers have a trust problem, and researchers claim a partial fix.
Generative reward models, or GRMs, are the AI judges used to train large language models, writing a natural-language critique alongside each preference call instead of just picking a winner. The catch is how they're trained: most GRMs are graded only on whether their final verdict was correct, so a critique with sloppy or missing reasoning can still get reinforced if it happens to land on the right answer. A new framework called EnGRICH addresses that by pairing the GRM with a second model, MetaCritic, trained on a small set of human-written critiques to build a custom rubric for each response and score whether the GRM's reasoning actually holds up. Those human-grounded standards are then extended to the much larger pool of preference data that has no human critique attached, and the approach improved results across seven reward-model benchmarks.
Reward models are the referees for the reinforcement-learning step that shapes chatbot behavior, so a referee that's right for the wrong reasons quietly bakes bad habits into every model trained against it. Human critique data is expensive and scarce, and most prior work either skips it or flattens it into a single number, throwing away exactly the detail that makes it useful. EnGRICH's contribution is squeezing more mileage out of that scarce resource by generalizing its lessons to outcome-only data instead of demanding more of it.
It's a sensible incremental fix, not a breakthrough, and a reminder that teaching AI to grade AI is still more alchemy than science.