[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-jailbreak-framework-cracks-ai-models-in-under-3-tries":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},6033,"new-jailbreak-framework-cracks-ai-models-in-under-3-tries","New Jailbreak Framework Cracks AI Models in Under 3 Tries","A new red-teaming method scripts AI personas to break six frontier models in just 2.46 messages, exposing a shared weak spot.","A new AI safety framework can talk six major chatbots into breaking their own rules in fewer than three messages, on average.\n\nResearchers built a system called BLUEPRINT that separates two things: a menu of 18 psychology-based persuasion tactics, like legitimacy appeals and gain framing, and a module called WORLDVIEWSIM that sets up a fictional scenario before the persuasion even starts. A search algorithm called Monte Carlo Tree Search tries different combinations of those tactics across a four-turn conversation, hunting for the sequence that gets a model to comply. Tested against six frontier models, both open-weight and proprietary, the method hit near-ceiling success rates while needing an average of just 2.46 messages per model. Even models that initially refused shared one common failure mode: they caved once the request was reframed as something concrete and immediately actionable.\n\nThe finding narrows down what actually breaks AI guardrails, and it isn't abstract jailbreak prompts - it's how believable and executable a request sounds once a model is deep in a conversation. That's a different threat model than the single-shot prompt injections most safety filters are built to catch, and it suggests defenses that only scan for harmful keywords are looking in the wrong place.\n\nEvery study on how easily AI breaks doubles as marketing for the next AI safety product, so treat the 2.46-message figure as a lab result, not a real-world crime rate - but the core claim, that concrete framing beats brute force, tracks with what red-teamers have argued for years.","[\"ai safety\",\"jailbreaking\",\"llm security\",\"red teaming\"]","2026-09-03T04:00:00.000Z","2026-09-03T06:07:56.175Z","2026-09-03T06:08:07.527Z","published",null,[],"ai",[26,27,28,29],"ai safety","jailbreaking","llm security","red teaming",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.02414",0,{"sections":36},[37,41,46,51,56,61,66,71,76,81,86,91,96,101],{"name":38,"slug":24,"count":39,"latest_published_at":40},"AI",3385,"2026-09-04T22:17:36.000Z",{"name":42,"slug":43,"count":44,"latest_published_at":45},"Security","security",565,"2026-09-05T00:03:08.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":50},"Policy","policy",300,"2026-09-04T22:18:34.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Hardware","hardware",152,"2026-09-03T09:26:48.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",97,"2026-09-04T15:29:18.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Science","science",96,"2026-09-03T22:30:00.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",54,"2026-09-04T23:36:14.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",37,"2026-09-04T20:22:41.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]