[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-fix-ai-overconfidence-without-new-training-data":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},10005,"researchers-fix-ai-overconfidence-without-new-training-data","Researchers Fix AI Overconfidence Without New Training Data","A new test-time method cuts LLM calibration error by an average of 70.8 percent across benchmarks without requiring any labeled training data.","Researchers built a way for AI models to grade their own confidence at test time, no answer key required.\n\nThe method, called Test-Time Calibration Learning (TTCL), comes from a paper posted to arXiv on October 5. Large language models often sound just as sure when they're wrong as when they're right, which is a problem if anyone downstream is trusting that confidence. TTCL fixes this by generating multiple responses to the same unlabeled question and using the agreement or disagreement between those responses as a stand-in for ground truth. Across eight benchmarks spanning math reasoning and factual question answering, it lifted base model accuracy by an average of 40.13 percent and cut calibration error, measured by ECE, by an average of 70.8 percent.\n\nMost calibration fixes to date have relied on reinforcement learning against labeled correctness data, which is exactly what's missing once a model is out in the wild facing new tasks. TTCL's label-free approach also helped models that were already calibrated but drifted after a domain shift: moving from math training to factual QA, it still produced a 20.35 percent accuracy gain and a 53.83 percent ECE reduction. That second result matters more than the headline numbers, since domain shift is the normal condition for deployed models, not the exception.\n\nThe catch is that self-supervision built from a model's own outputs can only catch disagreements the model is capable of noticing; a model that is confidently and consistently wrong across every sampled response would sail right through this check unflagged.","[\"ai\",\"llm calibration\",\"machine learning research\",\"test-time learning\"]","2026-10-05T04:00:00.000Z","2026-10-05T17:48:18.020Z","2026-10-05T17:48:23.366Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek says calibration error is cut 'by up to 70 percent' but the body reports 70.80% as an average across benchmarks, not a maximum\u002Frange — fix the dek to say 'an average of 70.8 percent' so it doesn't contradict the body's own framing.","resolved","ai",[30,32,33,34],"llm calibration","machine learning research","test-time learning",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.02695",0,{"sections":41},[42,45,49,54,59,64,68,73,77,82,87,92,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",6233,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",868,{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",444,"2026-10-03T15:02:01.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",323,"2026-10-04T13:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",204,"2026-10-03T14:50:50.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",177,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",158,"2026-10-03T03:21:12.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":18},"Dev Tools","dev-tools",98,{"name":78,"slug":79,"count":80,"latest_published_at":81},"Software","software",97,"2026-10-04T10:00:00.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",92,"2026-10-04T14:36:25.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",51,"2026-10-05T02:35:01.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",32,"2026-10-02T18:00:00.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]