A new scoring method asks AI models to grade each other's answers by fighting it out, not just rating them on a scale.
Researchers built MatrixReward, a reward mechanism for training AI on open-ended questions, the kind without one correct answer. Instead of grading each response against a rubric in isolation, it compares every pair of sampled answers under each rubric criterion and builds a win-rate matrix from the results. The spread and overlap within that matrix tell the system which rubrics are actually distinguishing good answers from bad ones in the current batch, so it weights those rubrics higher. Each answer is then scored by how close it lands to an ideal profile built from those weighted rubrics, and in tests on Qwen3-8B across four open-ended benchmarks, the method averaged 63.02, about 2 percent above the strongest existing baseline.
Training AI to write good essays, explanations, or open-ended answers is hard because there is no ground truth to check against, only judgment calls stacked on judgment calls. Most current approaches average several rubric scores into a single number, which tends to flatten exactly the differences that separate a mediocre answer from a sharp one. By weighting rubrics based on how much they actually differentiate the current set of answers, rather than treating every criterion as equally useful, MatrixReward points toward reward models that get smarter through better math rather than bigger models or more human labels.
A 2 percent gain on one 8-billion-parameter model is not a dramatic swing, but reward modeling probably needs more of these incremental, well-reasoned fixes than another benchmark headline.