[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-cant-prove-ai-resists-persuasion-study-finds":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},9453,"researchers-cant-prove-ai-resists-persuasion-study-finds","Researchers Can't Prove AI Resists Persuasion, Study Finds","A new study tests seven language models against a persuasive adversary in a deception game, then finds the results can't separate robustness from a coin flip.","A new study finds there's no reliable way to prove whether language models actually resist a persuasive liar, or just default to the most eye-catching answer.\n\nResearchers built a forced-choice task based on the party game Deception: Murder in Hong Kong, then added an adversary that knows the correct answer, watches the signal, and argues for the best wrong answer using a persuasion budget called beta. Across a 200,000-item pool, they calculated the adversary-robust optimal signal, the one that best survives a smart adversary, and compared it to a simpler salience heuristic from earlier work. The two targets point to the same answer on all but 2,748 items, and on the 108 confirmatory items used for testing, they overlap completely. Seven language models were then run through two adversary framings, and their chosen answers flipped on anywhere from 30 to 77 of those 108 items.\n\nThe flips look like evidence that persuasion pressure changes model behavior, and 18.2 percent of the full 200,000-item pool does have an optimum that shifts once the adversary's budget rises. But because the robust answer and the salience answer are identical on the test set used, there is no way to know if a model shifted because it is reasoning about the adversary or because it is leaning on a shortcut that happens to look the same. That is a measurement problem, not a minor caveat, for anyone building or buying AI systems meant to withstand manipulation.\n\nBefore trusting any 'adversary-robust' benchmark, check whether its answers are even distinguishable from a simpler guess, because this paper shows plenty aren't.","[\"ai\",\"adversarial robustness\",\"language models\",\"deception games\"]","2026-10-02T04:00:00.000Z","2026-10-02T19:46:01.537Z","2026-10-02T19:46:07.286Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the missing apostrophes in the headline and dek ('Cant' should be 'Can't') before publishing — the facts and figures otherwise check out exactly against the source.","resolved","ai",[30,32,33,34],"adversarial robustness","language models","deception games",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00233",0,{"sections":41},[42,45,49,54,59,64,69,74,79,84,89,94,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5765,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",831,{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",437,"2026-10-01T18:10:00.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",198,"2026-10-01T17:38:48.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Science","science",168,"2026-10-01T18:35:55.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]