AI/ reinforcement-learning · llms · ai-research · multimodal-ai

New Reward Method Lets AI Grade Its Own Reasoning Steps

A new reinforcement learning method scores reasoning steps using a model's own confidence, beating standard reward training without human annotations.

A new reinforcement learning technique scores each step of a language model's reasoning, not just the final answer, using nothing but the model's own confidence in its own outputs.

The method, called Stepwise Marginal Information Gain, comes from a paper posted to arXiv. Most reinforcement learning setups for large language models only reward a correct final answer, which makes it impossible to tell which steps in a chain of reasoning actually helped. The researchers instead track how much each reasoning step raises the model's own likelihood of producing the correct answer, using a watermark that only credits new highs so the model cannot loop through junk steps for extra reward. For vision-language models, they add a check that discounts reward when an answer looks correct without ever needing the image, aimed at stopping models from guessing off word patterns instead of what is actually in the picture.

On eight benchmarks, the technique beat standard outcome-only reward training in every comparison, including a 12.6-point jump on the MathVerse benchmark and a 12.9-point edge over an external process reward model that needed 16 samples per question, using only one. That matters because process reward models usually require expensive human-labeled step annotations or extra models running at inference time; this approach needs neither.

Whether the gains hold up outside curated benchmarks is the open question, as it usually is with self-reported arXiv results that have not gone through peer review.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →