[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-method-lets-local-ai-decide-when-to-ask-for-help":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8540,"new-method-lets-local-ai-decide-when-to-ask-for-help","New Method Lets Local AI Decide When to Ask for Help","HARISSA lets small on-device models judge their own answers, deciding when to reason harder and when to defer to a human instead of risking a wrong answer.","Local language models are getting better at knowing what they don't know.\n\nA new method called HARISSA teaches small, on-device language models to judge their own answers before committing to them. Researchers fine-tune a model so two of its internal states, one computed before it generates anything and one taken after it finishes an answer, each predict whether that answer will be correct. The system uses those predictions to decide, query by query, whether to spend extra computation on harder reasoning or to just defer the question to a human instead of guessing. Tested on a single device running one model, HARISSA matched chain-of-thought reasoning within one accuracy point while running 2.7 times faster; tested on a server hosting four sizes of the same model, it beat established cascading methods like FrugalGPT and Self-REF at matching latency, and left fewer wrong answers on the table than standard confidence scores in five of six task setups.\n\nThe core problem here isn't new: local models are private, cheap, and fast, but they are also small and less capable than anything running on a server. The usual fix, punting hard questions to the cloud, erases the privacy and cost benefits that made local deployment appealing in the first place. HARISSA's pitch is that a model can make both the reasoning-effort call and the answer-or-defer call using signals it already has internally, without leaving the device.\n\nThat's a real result, but it comes from arXiv benchmarks, not a shipped product, and the safety gains depend on how well those internal signals generalize past the tasks tested here. Still, as phone and laptop makers keep pushing on-device assistants, techniques like this are what will decide whether local AI stays a genuinely private option or just a worse version of the cloud model it's meant to replace.","[\"ai\",\"on-device ai\",\"inference\",\"research\"]","2026-09-30T04:00:00.000Z","2026-09-30T10:01:26.832Z","2026-09-30T10:01:29.722Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The closing paragraph is caveat-only and stops abruptly — end with a sentence that ties the caveats back to what this means for readers\u002Fthe industry rather than trailing off on unmonitored deferral.","resolved","ai",[30,32,33,34],"on-device ai","inference","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38006",0,{"sections":41},[42,45,49,53,58,63,68,73,78,82,87,92,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5104,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",785,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",417,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]