[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-models-still-fail-at-verifying-security-protocols-study-finds":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},6981,"ai-models-still-fail-at-verifying-security-protocols-study-finds","AI Models Still Fail at Verifying Security Protocols, Study Finds","A new benchmark pits GPT and DeepSeek against formal verification tools on 388 security goals, and both models still miss attacks formal tools would catch.","A new study put GPT and DeepSeek up against the tools security researchers actually trust, and the chatbots still don't cut it.\n\nResearchers tested both models, in chat and reasoning modes, on 130 obfuscated security protocols covering 388 verification goals, then scored the results against established formal verification tools ProVerif and OFMC. In chat mode, both models flagged problems aggressively but wrongly: GPT hit 72.7% recall at just 27.3% precision, and DeepSeek scored 69.3% recall at 27.2% precision, meaning most flagged goals were false alarms. Switching on reasoning mode flipped that trade-off: GPT's precision rose to 66.5% (recall fell to 54.5%), and DeepSeek's precision rose to 45.4% (recall 57.3%). Worked out against the underlying counts, GPT's reasoning mode edges past a trivial always-secure guess, which scores 77.1% accuracy on this lopsided dataset of 89 vulnerable and 299 secure goals; every other model and mode fell short of that low bar.\n\nThe split that should worry security teams is by category. Authentication goals are where reasoning models catch well under half of real attacks, while chat mode buries the hits it does find under a pile of false positives. Confidentiality checks fared far better, with reasoning mode F1 scores reaching 95.7%, meaning these models are better at spotting exposed secrets than at catching someone impersonating another party.\n\nResults also weren't stable. Rerunning the same prompt three times produced identical verdicts on just 61.6% to 89.7% of goals depending on model and mode, and the models' own confidence scores didn't track whether they were actually right. The researchers' own verdict: at best, an LLM might work as a pre-screening filter before a real formal verification tool, not a replacement for one.","[\"ai\",\"security\",\"benchmarks\",\"research\"]","2026-09-18T04:00:00.000Z","2026-09-19T01:38:16.732Z","2026-09-19T01:38:28.657Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"publisher-r1","publisher",1,"The article states GPT's reasoning mode (66.5% precision, 54.5% recall) beat the 77.1% always-safe baseline, but 66.5% precision at 54.5% recall implies an accuracy well below 77.1%, so that claim is internally inconsistent.","resolved","ai",[30,32,33,34],"security","benchmarks","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.20712",0,{"sections":41},[42,45,49,54,59,63,67,72,76,81,86,91,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",4083,{"name":46,"slug":32,"count":47,"latest_published_at":48},"Security",663,"2026-09-18T14:00:14.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",339,"2026-09-17T12:00:00.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":18},"Hardware","hardware",155,{"name":64,"slug":65,"count":66,"latest_published_at":18},"Science","science",125,{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":18},"Dev Tools","dev-tools",78,{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]