[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-study-finds-ai-vision-models-stay-confident-when-wrong":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},6528,"new-study-finds-ai-vision-models-stay-confident-when-wrong","New Study Finds AI Vision Models Stay Confident When Wrong","An unreviewed arXiv preprint finds vision-language models voice equal confidence in correct and garbled reasoning, exposing a blind spot in AI self-checks.","AI systems that say \"let me double-check\" can still land on the wrong answer and sound just as sure of themselves as if they had nailed it.\n\nA new preprint, \"The Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models\" (arXiv:2609.18453, https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.18453, posted September 17, 2026, and not yet peer-reviewed), tested how vision-language models voice confidence while working through multi-step reasoning. The researchers found that a model's stated confidence barely shifts based on what its reasoning trajectory actually contains, whether the steps are sound, muddled, or edited to strip out key information. They checked this three ways: swapping content within the reasoning chain, masking individual tokens, and analyzing the model's own hesitation phrases, like \"wait, I should recheck.\" Standard calibration metrics such as ECE and AUROC missed the problem entirely, so the authors built a new benchmark, TGS-Bench, using paired good and bad reasoning trajectories across 10 tasks to measure it directly.\n\nThat gap matters because confidence scores are increasingly used as a shortcut for trust, telling a system, a reviewer, or a user when an AI answer is safe to accept without a second look. If confidence tracks nothing about the reasoning that produced it, those scores are closer to a costume than a measurement, particularly in tasks that mix vision and multi-step logic. The paper's more unsettling finding: calibration training, the standard fix for overconfident models, can make this disconnect worse rather than better.\n\nA model that says it is rechecking its work and then confidently repeats the same wrong answer is not demonstrating judgment. It is demonstrating a script. Anyone bolting confidence thresholds onto an automated pipeline should read the fine print before trusting the number.","[\"ai\",\"research\",\"vision-language-models\",\"model-evaluation\"]","2026-09-17T04:00:00.000Z","2026-09-17T22:38:10.669Z","2026-09-17T22:38:22.583Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Attribute the study properly in the body — name the paper (\"The Mirage of Calibrated Confidence...\"), note it's an unreviewed arXiv preprint, and include the arXiv ID\u002Flink, since right now it's sourced only to unnamed 'researchers' with no institution or citation given.","resolved","ai",[30,32,33,34],"research","vision-language-models","model-evaluation",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.18453",0,{"sections":41},[42,46,50,55,60,64,68,73,78,82,87,92,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",3852,"2026-09-17T08:27:09.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",648,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":18},"Hardware","hardware",154,{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",114,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",73,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]