[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-on-device-ai-models-get-a-reality-check-for-entity-extraction":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},9581,"on-device-ai-models-get-a-reality-check-for-entity-extraction","On-Device AI Models Get a Reality Check for Entity Extraction","A new study pits a tiny spaCy tagger and GLiNER encoders against local LLMs for named entity recognition, finding smaller models win on speed and reliability.","A new benchmark says the smallest AI model is often the right one for pulling names, places, and organizations out of text on your own device.\n\nResearchers tested nine on-device systems across three families: one classical spaCy tagger, three GLiNER encoder models (166 to 460 million parameters), and five generative large language models run locally - Qwen3 at 0.6B, 1.7B, and 4B parameters, plus DeepSeek-R1 at 1.5B and 8B. They ran all nine against three datasets and measured accuracy alongside two things most leaderboards skip: latency and output validity, meaning whether an answer is even well-formed. Because one dataset lacked verified ground truth, the team built silver labels from a panel of LLM judges, then checked that panel's work against a full human re-annotation. The gold standard mattered: switching from LLM-generated labels to human-verified ones made every encoder model look better and every generative model look worse.\n\nOn raw accuracy, a 4B-parameter instruct model can beat everything else on clean newswire text. But deployability is a different contest, and GLiNER's small encoders matched or nearly matched that performance at one-ninth to one-twenty-fourth the size, with millisecond-to-second latency and zero malformed outputs. The smallest generative models, by contrast, produced invalid output up to 27% of the time on long inputs - a problem that scale fixed and a bigger output budget did not.\n\nGLiNER's confidence scores are also overconfident, not just inaccurate - calibration error as high as 0.47, roughly halved by standard temperature scaling - a reminder that a model saying it's sure on-device is not the same as a model being right.","[\"ai\",\"on-device-ai\",\"named-entity-recognition\",\"benchmarking\"]","2026-10-02T04:00:00.000Z","2026-10-03T01:35:26.015Z","2026-10-03T01:35:30.906Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the count mismatch: the lead says nine AI systems, but the breakdown (1 spaCy + 5 GLiNER + 4 generative) adds up to ten — recheck against the source, which lists five generative models (Qwen3-0.6B\u002F1.7B\u002F4B plus DeepSeek-R1-1.5B\u002F8B), so the GLiNER count is likely off; make the breakdown sum to nine.","resolved","ai",[30,32,33,34],"on-device-ai","named-entity-recognition","benchmarking",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00007",0,{"sections":41},[42,45,49,53,58,62,66,71,76,81,86,91,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5896,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",837,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",438,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":18},"Hardware","hardware",199,{"name":63,"slug":64,"count":65,"latest_published_at":18},"Science","science",171,{"name":67,"slug":68,"count":69,"latest_published_at":70},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]