[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-hallucination-detectors-have-blind-spots-new-study-finds":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8560,"hallucination-detectors-have-blind-spots-new-study-finds","Hallucination Detectors Have Blind Spots, New Study Finds","A new study finds popular hallucination detectors work well on some AI errors and almost not at all on others, with the gap varying wildly by model.","Hallucination detectors don't fail evenly - and a new arXiv paper puts a number on just how uneven.\n\nResearchers tested sampling-based consistency detection, the common technique of asking a model the same question multiple times and checking whether its answers agree, across four language models and three factual question-answering datasets. They split hallucinations into two groups: \"Ghost\" cases, where the model's repeated answers mostly agree even though the answer is wrong, and \"Flickering\" cases, where answers disagree a lot. The gap between how detectable these two groups are came out to 0.35-0.46 AUC (area under the curve), a 0-to-1 score for how well a detector separates hallucinated answers from correct ones, where 0.5 is a coin flip and 1.0 is perfect. Because the math used to define the groups was itself correlated with the math used to measure the gap, the team re-ran the test after locking in the group assignments, and the same asymmetry showed up in independent measures of answer wording and meaning, holding up across all 12 model-and-dataset combinations and surviving a stricter check on two additional model families.\n\nThat means the same detector can look reliable on average while quietly missing an entire class of errors on a given model. The share of hallucinations landing in the harder-to-catch group ranged from 16% to 77% depending on which model was tested, and the same prompt often flipped between easy and hard depending on the model answering it.\n\nAny vendor citing a single hallucination-detection score is telling you less than it sounds like - the interesting number is the one they're not showing you: how that score splits by model and by case.","[\"hallucination detection\",\"ai research\",\"llm evaluation\",\"arxiv\"]","2026-09-30T04:00:00.000Z","2026-09-30T11:16:46.423Z","2026-09-30T11:16:50.204Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The source measures the 0.35-0.46 gap and the 16%-77% blind-spot range in AUC (area under curve, a detector's ability to distinguish hallucinations from correct answers), not 'accuracy' as the draft states — correct the metric label and briefly explain what AUC means for readers.","resolved","ai",[32,33,34,35],"hallucination detection","ai research","llm evaluation","arxiv",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.35860",0,{"sections":42},[43,46,50,54,59,64,69,74,79,83,88,93,98,103],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5105,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",785,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",417,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]