[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-new-llm-defense-rewrites-prompts-to-expose-hidden-jailbreaks":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},6224,"a-new-llm-defense-rewrites-prompts-to-expose-hidden-jailbreaks","A New LLM Defense Rewrites Prompts to Expose Hidden Jailbreaks","A new black-box defense rewrites suspicious prompts to expose hidden intent before scoring them for risk, catching disguised jailbreaks other filters miss.","Researchers have proposed a filter that catches AI jailbreaks by rewriting suspicious prompts to strip away their disguise before judging them.\n\nThe system, called SRD-GUARD, works without needing access to a model's internal code. Given a prompt, it generates five reworded versions that strip away fictional framing, role-play setups, or other contextual disguise while keeping the underlying request intact, then runs the original and all five rewrites through multiple independent AI safety scorers. A decision module weighs the absolute risk score and how sharply it jumps between the original and its rewrites, then decides whether to block, allow, or flag the request. Tested against three known jailbreak techniques (UNIATTACK, CIPHER, and DeepInception) on an uncensored Llama 3 variant and DeepSeek V4 Flash, it averaged defense success rates of 91.44% and 100% while wrongly rejecting legitimate requests only 8% to 12% of the time, a better balance than the compared baselines.\n\nJailbreaks that dress up harmful requests as screenplay scenes, research hypotheticals, or role-play have been a persistent hole in AI guardrails, and most existing defenses either miss the disguise or refuse too many innocent prompts. Stripping the packaging away before judging the content is a sensible fix for that specific failure mode, and the reported over-refusal rates are notably lower than typical filter trade-offs. But running five rewrites through multiple scorers for every single prompt means several times more inference calls than a basic filter, a cost the paper never quantifies.\n\nThe numbers come from a single, not-yet-peer-reviewed arXiv paper tested on a deliberately uncensored model and released via an anonymized code repository, so this reads as a promising lab result rather than something ready for a production chatbot.","[\"ai-safety\",\"jailbreak-defense\",\"llm-security\",\"research\"]","2026-09-10T04:00:00.000Z","2026-09-10T07:36:58.904Z","2026-09-10T07:37:10.824Z","published",null,[],"security",[26,27,28,29],"ai-safety","jailbreak-defense","llm-security","research",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.06540",0,{"sections":36},[37,42,45,50,55,60,65,69,74,79,84,89,94,99],{"name":38,"slug":39,"count":40,"latest_published_at":41},"AI","ai",3480,"2026-09-11T04:00:00.000Z",{"name":43,"slug":24,"count":44,"latest_published_at":41},"Security",629,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",336,"2026-09-11T00:56:21.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Hardware","hardware",153,"2026-09-09T15:12:32.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":41},"Science","science",98,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]