Local language models are getting better at knowing what they don't know.
A new method called HARISSA teaches small, on-device language models to judge their own answers before committing to them. Researchers fine-tune a model so two of its internal states, one computed before it generates anything and one taken after it finishes an answer, each predict whether that answer will be correct. The system uses those predictions to decide, query by query, whether to spend extra computation on harder reasoning or to just defer the question to a human instead of guessing. Tested on a single device running one model, HARISSA matched chain-of-thought reasoning within one accuracy point while running 2.7 times faster; tested on a server hosting four sizes of the same model, it beat established cascading methods like FrugalGPT and Self-REF at matching latency, and left fewer wrong answers on the table than standard confidence scores in five of six task setups.
The core problem here isn't new: local models are private, cheap, and fast, but they are also small and less capable than anything running on a server. The usual fix, punting hard questions to the cloud, erases the privacy and cost benefits that made local deployment appealing in the first place. HARISSA's pitch is that a model can make both the reasoning-effort call and the answer-or-defer call using signals it already has internally, without leaving the device.
That's a real result, but it comes from arXiv benchmarks, not a shipped product, and the safety gains depend on how well those internal signals generalize past the tasks tested here. Still, as phone and laptop makers keep pushing on-device assistants, techniques like this are what will decide whether local AI stays a genuinely private option or just a worse version of the cloud model it's meant to replace.