[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-tool-maps-how-chatbots-break-under-multi-turn-pressure":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},5977,"new-tool-maps-how-chatbots-break-under-multi-turn-pressure","New Tool Maps How Chatbots Break Under Multi-Turn Pressure","Researchers built an evolutionary search system that maps how AI chatbots fail under multi-turn manipulation, and Claude proved hardest to crack.","A new red-teaming framework called EvoFlint shows that even Claude Sonnet 4.6 still folds to a patient enough conversation.\n\nEvoFlint treats multi-turn jailbreaking as a search problem rather than a generation problem. Instead of grinding out one-off prompts, it evolves phased conversation plans through LLM-driven mutation and crossover, scoring each on a mix of attack success rate and peak severity so near-misses still shape the next generation. A risk-indexed archive runs novelty search to keep the strategy pool diverse without locking into a fixed style taxonomy, and a shared memory feeds insights about each target model back into new attempts. Tested on the HarmBench-test benchmark, EvoFlint reached a 35.8% attack success rate against Claude Sonnet 4.6, 59.7% against GPT-5.4, 94.3% against Qwen3-32B, and 98.7% against the older GPT-4o.\n\nThat spread is the real finding. Most automated red-teaming produces a pile of prompts that get patched and forgotten; EvoFlint instead produces a persistent, per-category map of which harms a model's safety training actually covers, which is more useful to defenders than a single pass\u002Ffail score. The gap between a 35.8% break rate on a frontier model and a 94.3% break rate on an open-weight one also says something about where safety tuning budgets are, and aren't, being spent.\n\nWorth remembering that today's hardest target is tomorrow's baseline: the paper's own GPT-4o numbers show a model that looked safe on release now failing nearly every attack.","[\"ai-safety\",\"red-teaming\",\"llm-security\",\"jailbreaking\"]","2026-09-02T04:00:00.000Z","2026-09-02T06:44:06.404Z","2026-09-02T06:44:18.297Z","published",null,[],"security",[26,27,28,29],"ai-safety","red-teaming","llm-security","jailbreaking",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.00487",0,{"sections":36},[37,42,46,51,56,61,66,71,76,81,86,91,96,101],{"name":38,"slug":39,"count":40,"latest_published_at":41},"AI","ai",3385,"2026-09-04T22:17:36.000Z",{"name":43,"slug":24,"count":44,"latest_published_at":45},"Security",565,"2026-09-05T00:03:08.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":50},"Policy","policy",300,"2026-09-04T22:18:34.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Hardware","hardware",152,"2026-09-03T09:26:48.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",97,"2026-09-04T15:29:18.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Science","science",96,"2026-09-03T22:30:00.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",54,"2026-09-04T23:36:14.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",37,"2026-09-04T20:22:41.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]