[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-miscalibrated-threshold-erases-correct-answers-in-language-models":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},6133,"a-miscalibrated-threshold-erases-correct-answers-in-language-models","A Miscalibrated Threshold Erases Correct Answers in Language Models","A new study finds a tiny language model has the right verdict buried in its output scores, but a single miscalibrated threshold flips it to yes every time.","A small language model answered every logic-verification question with \"yes\" - even when the answer was wrong - but the correct verdict was sitting untouched in its output scores the whole time.\n\nResearchers tested a 0.6-billion-parameter model on 1,200 logical conclusions, half valid and half broken by a single edited word. The model said yes to all of them, landing at 50% behavioral accuracy. Linear probes on the model's hidden states, though, correctly read the right verdict 96% of the time, and that signal held up even on logic structures the probes had never seen. Tracing the failure further, the researchers found the correct answer survived all the way to the model's raw output scores (89% AUC) - it just never crossed the decision threshold that turns those scores into a spoken answer, because that threshold was miscalibrated by more than four and a half standard deviations.\n\nThat reframes a chunk of \"hallucination\" and \"models don't know what they don't know\" research: sometimes the knowledge is there and correctly computed, but a single broken dial between a model's internal math and its printed output throws it away. The fix was cheap - recalibrating that one threshold pushed the small model's accuracy from 50% to 81%, and a calibrated decoding trick recovered 94% on an 8-billion-parameter model, without retraining anything.\n\nThat also explains a stranger finding in the paper: an 8B model did worse on this task than its own 4B sibling, not because it reasoned less well, but because its output dial was more broken - a reminder that bigger models aren't automatically a fix for a plumbing problem.","[\"ai-research\",\"llm-interpretability\",\"model-calibration\",\"arxiv\"]","2026-09-07T04:00:00.000Z","2026-09-07T06:39:53.694Z","2026-09-07T06:40:05.595Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the headline\u002Fdek framing — 'refuse to say' implies intentional withholding, but the source describes a miscalibrated decision threshold erasing a correct internal signal, not the model choosing not to answer; reword to reflect the actual mechanism (miscalibration causing wrong yes\u002Fno output) rather than anthropomorphized refusal.","resolved","ai",[32,33,34,35],"ai-research","llm-interpretability","model-calibration","arxiv",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.04582",0,{"sections":42},[43,47,52,57,62,67,72,77,82,87,92,97,102,107],{"name":44,"slug":30,"count":45,"latest_published_at":46},"AI",3411,"2026-09-07T10:06:25.000Z",{"name":48,"slug":49,"count":50,"latest_published_at":51},"Security","security",574,"2026-09-07T07:03:20.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Policy","policy",311,"2026-09-07T05:33:25.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":61},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":66},"Hardware","hardware",152,"2026-09-03T09:26:48.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",97,"2026-09-04T15:29:18.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Science","science",96,"2026-09-03T22:30:00.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Startups","startups",54,"2026-09-04T23:36:14.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"General","general",40,"2026-09-07T08:57:14.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]