AI/ reinforcement-learning · llm-training · ai-research

RL Training Trick Fixes a Blind Spot in LLM Fine Tuning

A new masking method called CARM catches canceling probability swings that an older filtering technique missed, lifting math and code benchmark scores.

Researchers have found a flaw in a popular technique for keeping AI training on track, and built a fix that boosts performance on math and coding tests.

When companies fine-tune large language models with reinforcement learning, the system that generates responses and the system that learns from them can drift out of sync. A common fix discards responses that drifted too far, measured by averaging how much each word's probability changed. The catch: that average can hide the problem. If some words became much more likely and others much less likely, the swings can cancel out mathematically, even though the underlying drift was large. The new method, called Cancellation-Aware Response Masking (CARM), fixes this by averaging the absolute size of the changes instead of letting positive and negative shifts offset each other.

This matters because off-policy drift is a known source of instability in RL-trained models, and a masking rule that looks clean on paper but quietly lets bad data through can waste expensive training runs. On AIME math benchmarks, CARM improved scores by up to 3.13 percentage points over the older geometric-mean approach, and lifted average code-generation accuracy by 2.88 points over the best baseline tested.

It is a small fix to a specific plumbing problem, not a new way of training models. But in a field where labs are pouring compute into RL post-training for reasoning, a free few points of accuracy from fixing a measurement bug is the kind of unglamorous win that tends to get quietly adopted rather than announced with fanfare.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →