[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ais-confident-wrong-answers-are-stable-not-just-fragile":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},5004,"ais-confident-wrong-answers-are-stable-not-just-fragile","AI's Confident Wrong Answers Are Stable, Not Just Fragile","New research finds some high-confidence AI errors resist small nudges just like correct answers, complicating simple fixes like self-critique prompting.","Some of AI's most confident wrong answers are not glitches. They are stable.\n\nA new arXiv paper studies \"stable miscalibration\" in large language models, cases where a confident wrong answer holds steady even when the input is slightly perturbed, rather than wobbling like a shaky guess. The researchers built two diagnostics: an audit score that flags domains with mismatched confidence and forced-answer overconfidence, and an internal probe that tracks how much a model's hidden states shift under perturbation. Testing across a multi-domain binary factual set, they found that self-critical prompting, where a model is asked to reconsider or abstain, consistently reduced that internal hidden-state sensitivity across layers in three open-weight models. But the audit score's link to actual decision-loss reduction was weaker than direct labeled baselines, and the paper found no clear evidence that these overconfident errors were internally shakier than correct answers.\n\nThat last point matters more than it sounds. Most efforts to fix AI hallucinations assume wrong answers are computationally fragile, that nudging the model should crack the error loose. This research suggests some wrong answers are just as internally settled as correct ones, meaning confidence, whether read by a human or measured inside the model, is a poor proxy for correctness.\n\nSelf-critique prompting still changed how settled a model's internal state looked. It just did not reliably change whether the model was right. Stability and truth, it turns out, are not the same thing, and this paper is among the first to measure the gap between them directly.","[\"ai\",\"llm calibration\",\"hallucinations\",\"ai research\"]","2026-08-17T04:00:00.000Z","2026-08-17T05:01:48.045Z","2026-08-17T05:01:59.879Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Remove or attribute the closing claim that self-critique prompting is already 'a go-to patch for hallucinations in production chatbots' — this specific claim about production deployment isn't corroborated anywhere in the supplied source material and should be cut or sourced separately.","resolved","ai",[30,32,33,34],"llm calibration","hallucinations","ai research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.13591",0,{"sections":41},[42,46,50,55,60,65,70,75,80,85,90,95,100,105],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":45},"Security","security",435,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]