Researchers just built a test to see whether AI can catch mistakes in doctors' notes - in two languages, not just English.
The benchmark, called MedRECT, pulls 663 error-laced passages from Japan's medical licensing exams and 458 more from the existing MEDEC dataset, then asks models to do three things: spot an error, identify which sentence has it, and fix it. The team ran 11 models, mixing proprietary and open-weight systems, through 17 different configurations, some with medical fine-tuning and some with extended reasoning turned on. Qwen3-32B did sharply better when reasoning was switched on, improving sentence-extraction accuracy by 24.5 percentage points on the Japanese set and 10.3 points on the English one. Several general-purpose reasoning models also beat all three medical-specialized models on detection and extraction.
That gap matters because hospitals shopping for AI scribes or chart-checking tools often assume a medical-tuned model will outperform a general one by default. This benchmark says that's not a safe bet, and that turning on a model's reasoning mode may do more than bolting on medical training data. It's also one of the few clinical-error datasets built outside English, which matters given how much deployed healthcare AI gets trained and graded on English text alone.
LoRA fine-tuning did close some of the gap on correction accuracy, so specialization isn't worthless - it's just not the whole story the medical-AI pitch decks tell.