A new training method claims to fix the OCR mistakes that slip through even after a model has already been fine-tuned to read documents.
Researchers built SP-DocReader, a self-play system that keeps using a model's own practice runs to find and fix remaining reading errors, rather than retraining it from scratch. It compares a model's generated text against the correct answer, lines up the matching words, and flags exactly where the two diverge. Those flagged spots get extra, direct training signal to push the model toward the right answer, while the rest of the model stays frozen. Tested on Qwen3-VL-4B, the technique cut character error rate by about 54 percent compared to a standard fine-tuning baseline, and improved document question answering accuracy by 3.7 points on the ANLS metric.
That's a meaningful jump for a field that usually chases accuracy by throwing more data or bigger models at a problem. This approach instead targets the specific failures that remain after normal training, which could make smaller, cheaper models usable for document-heavy tasks like processing scanned forms, invoices, and receipts without a full retrain.
The results come from the paper's own benchmarks and baseline comparisons, not an independent test, so how much of that gain survives contact with messy real-world scans, handwriting, or non-English documents remains to be seen.