A new training technique keeps AI models from quietly "fixing" text they can't quite read.
Researchers describe Gated and Attenuated on-policy Distillation (GAD-RL), a post-training method for vision-language models used in OCR. The problem: these models sometimes rewrite garbled or anomalous text in images into something that reads more naturally, which undermines transcription accuracy. The team found that guidance from a fixed "teacher" model becomes less useful as the student model improves, so GAD-RL adjusts how much it leans on that teacher based on the student's current performance - turning off distillation entirely once a response hits a task reward of at least 0.95, and gradually dialing it back before that point. On Qwen3.5-2B, GAD-RL scored 59.92% Micro Recall on the CHAOS-Bench benchmark, beating the GRPO and GRPO+OPD baselines by 8.45 and 4.43 percentage points, and posted an Overall score of 91.18 on OmniDocBench v1.6.
OCR is supposed to report what's actually in an image, not what sounds plausible. As more products stack language models on top of OCR pipelines - scanning IDs, receipts, medical forms - the habit of "autocorrecting" unusual text, like typos, redacted fields, or foreign characters, stops being a curiosity and starts being a liability. A method that explicitly tunes how much a model trusts its own fluency instincts versus the literal source material targets a narrow but real failure mode.
It's a reminder that a more fluent model isn't automatically a more faithful one - sometimes fluency is exactly what needs training out.