[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-vlms-know-when-theyre-guessing-but-answer-anyway":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},4957,"vlms-know-when-theyre-guessing-but-answer-anyway","VLMs Know When They're Guessing but Answer Anyway","A new benchmark shows vision-language models can internally detect when video evidence is inconclusive but still fail to admit uncertainty out loud.","AI models that analyze video often know when they're bluffing.\n\nResearchers built TRAPSBench, a benchmark of 1,404 paired physics videos where a small change makes the outcome impossible to determine from the footage alone. Testing 16 vision-language models across five families, the best model scored just 0.292 on a new metric called Penalized Epistemic Calibration Score, which rewards correct answers when the outcome is knowable and abstention when it isn't. The models mostly guessed instead of admitting they couldn't tell. But probing their internal states told a different story: linear probes could predict whether a video's outcome was actually answerable with up to 0.91 AUROC, and researchers could flip a single internal \"direction\" to switch abstention on or off.\n\nThat gap matters because it reframes a familiar failure. AI hallucination is usually described as models not knowing what they don't know. This paper suggests that's often wrong, at least for vision. The information needed to say \"I can't tell\" is sitting right there in the model's hidden layers. Something between that internal signal and the words coming out is dropping the ball.\n\nThe researchers also found models are far better at flagging uncertainty in text than in video, catching textual impossibility about four times more readily than missing visual evidence. That tracks with how these models are trained: heavily on text, with vision treated more as an add-on. If a self-driving system or a video-based safety tool inherits this pattern, it's not a knowledge problem, it's a plumbing problem, and one that output-level fixes might actually solve without retraining the whole model.","[\"vision-language models\",\"ai benchmarks\",\"hallucination\",\"model evaluation\"]","2026-08-14T04:00:00.000Z","2026-08-14T21:32:30.510Z","2026-08-14T21:32:42.410Z","published",null,[],"ai",[26,27,28,29],"vision-language models","ai benchmarks","hallucination","model evaluation",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.13167",0,{"sections":36},[37,41,45,50,55,60,65,70,75,80,85,90,95,100],{"name":38,"slug":24,"count":39,"latest_published_at":40},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":42,"slug":43,"count":44,"latest_published_at":40},"Security","security",435,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]