Researchers just showed that small, open-source AI models can do document-scanning work institutions have been handing to big commercial APIs.
A team evaluated eight open-source vision-language models, each with up to 7 billion parameters, on the job of turning scanned document images into clean, schema-compliant JSON. They tested across three university heritage collections containing historical handwriting, specialist vocabulary from fields like jewelry-making, prehistory, and architecture, and layouts that don't follow any standard template. The models were run zero-shot, with a handful of examples, and after fine-tuning, with results scored on character error rate, text-similarity, and structured-output accuracy. The researchers also tested whether hyperparameter tuning, classic image cleanup such as denoising and contrast adjustment, and multi-stage training each moved the needle independently.
This matters because heritage archives have real reasons to avoid sending scans to closed commercial VLMs: recurring API costs, data-autonomy concerns about historical records, and the energy footprint of hyperscale inference. A 7B-parameter model an institution can fine-tune and run on its own hardware sidesteps all three, provided it hits comparable accuracy - which this study suggests is achievable with the right preprocessing and training recipe.
The catch is in the details: a single checkpoint trained across all three collections can run anywhere, but the study found its accuracy shifts between datasets compared to training one model per collection. So "good enough, local, and free of API bills" is real, but it's not yet a plug-and-play swap for whatever commercial system archivists currently lean on.