A new benchmark finds the biggest energy savings in local document-reading AI come from batching requests, not from swapping in a smaller model.
The researchers tested small, on-premise AI models (both text-only and vision-language, capped at 8 billion parameters) on two very different document sets: Kleister-NDA contracts, which read close to plain text, and VRDU forms, which are dense with layout. They tracked both accuracy and energy use across preprocessing method, model family, and inference settings like batching and quantization. Batching turned out to be the biggest lever, cutting energy per page by 38-85% without hurting accuracy, while FP8 quantization saved 27-32% for one-off requests but only 9-19% (under 1 mWh per page) once batching was already in play. Neural OCR, meanwhile, burned 17 times more energy per page than classical OCR and never landed on the accuracy-efficiency frontier.
That matters because privacy-sensitive documents (medical records, contracts, government forms) often can't touch a cloud API, so efficiency has to be engineered into the local setup itself. The study's bigger finding is that there's no universal winner: vision-language models pay off on messy, layout-heavy forms, but for near-plain-text documents, a small text model paired with a cheap OCR parser beats every vision-language option on both accuracy and energy.
It's a useful reminder that a fancier model isn't always the answer: a batching setting and an old-school OCR parser did more for the energy bill here than any model swap.