A new benchmark says the smallest AI model is often the right one for pulling names, places, and organizations out of text on your own device.
Researchers tested nine on-device systems across three families: one classical spaCy tagger, three GLiNER encoder models (166 to 460 million parameters), and five generative large language models run locally - Qwen3 at 0.6B, 1.7B, and 4B parameters, plus DeepSeek-R1 at 1.5B and 8B. They ran all nine against three datasets and measured accuracy alongside two things most leaderboards skip: latency and output validity, meaning whether an answer is even well-formed. Because one dataset lacked verified ground truth, the team built silver labels from a panel of LLM judges, then checked that panel's work against a full human re-annotation. The gold standard mattered: switching from LLM-generated labels to human-verified ones made every encoder model look better and every generative model look worse.
On raw accuracy, a 4B-parameter instruct model can beat everything else on clean newswire text. But deployability is a different contest, and GLiNER's small encoders matched or nearly matched that performance at one-ninth to one-twenty-fourth the size, with millisecond-to-second latency and zero malformed outputs. The smallest generative models, by contrast, produced invalid output up to 27% of the time on long inputs - a problem that scale fixed and a bigger output budget did not.
GLiNER's confidence scores are also overconfident, not just inaccurate - calibration error as high as 0.47, roughly halved by standard temperature scaling - a reminder that a model saying it's sure on-device is not the same as a model being right.