[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-research-flags-words-that-trick-ai-into-unsafe-replies":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},7565,"new-research-flags-words-that-trick-ai-into-unsafe-replies","New Research Flags Words That Trick AI Into Unsafe Replies","A new attribution model spots which turns and words nudge multi-turn AI conversations toward unsafe outputs, catching manipulation keyword filters miss.","A new AI safety model does not just flag a risky chatbot conversation - it points to the specific turn and words that pushed it there.\n\nResearchers built a multi-turn dataset of 1,762 conversations, mixing adversarial exchanges, \"benign twin\" conversations that look similar but aren't harmful, and benign chats stuffed with high-risk vocabulary. They trained a lightweight hierarchical model to both flag unsafe conversations and attribute the violation to specific user turns and token spans. The model hit an F1 score of 0.988 on detection, and stripping out just the top 15% of flagged tokens cut its confidence that a conversation was adversarial by 51.1%. Independent human reviewers largely agreed with the model's picks: its top five attributed turns included a human-identified evidence turn in 84.5% of adversarial cases.\n\nThe real story is what this fixes: keyword-based filters. A simple keyword baseline threw false positives on 37.3% of borderline-benign chats and a startling 94.7% of benign chats that just happened to use high-risk vocabulary - think security researchers discussing malware, or nurses discussing overdoses. This new model kept both false-positive rates under 1%, which matters more than the headline detection score to anyone who has had a legitimate conversation blocked by an overzealous filter.\n\nIt is still a research paper, not a feature sitting in front of a chatbot today. But as AI systems get pushed into agentic roles where a slow-drip manipulation can trigger a real-world action, turn-level attribution - not just a yes-or-no safety verdict - looks like the direction guardrails need to go.","[\"ai-safety\",\"llm-guardrails\",\"multi-turn-attacks\",\"research\"]","2026-09-24T04:00:00.000Z","2026-09-24T06:55:04.666Z","2026-09-24T06:55:09.853Z","published",null,[],"ai",[26,27,28,29],"ai-safety","llm-guardrails","multi-turn-attacks","research",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.27773",0,{"sections":36},[37,40,44,49,54,59,64,69,74,79,84,89,94,99],{"name":38,"slug":24,"count":39,"latest_published_at":18},"AI",4424,{"name":41,"slug":42,"count":43,"latest_published_at":18},"Security","security",724,{"name":45,"slug":46,"count":47,"latest_published_at":48},"Policy","policy",380,"2026-09-23T22:53:43.000Z",{"name":50,"slug":51,"count":52,"latest_published_at":53},"Deals","deals",227,"2026-09-24T11:08:33.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Hardware","hardware",174,"2026-09-24T10:10:29.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Science","science",136,"2026-09-24T09:00:00.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Consumer Tech","consumer-tech",116,"2026-09-24T00:51:49.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Software","software",85,"2026-09-23T20:00:00.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Dev Tools","dev-tools",79,"2026-09-22T22:21:13.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Startups","startups",66,"2026-09-23T17:28:38.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Gaming","gaming",45,"2026-09-22T15:35:06.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"General","general",43,"2026-09-21T23:48:56.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Reviews","reviews",27,"2026-09-22T13:00:00.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]