[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-build-automated-jailbreak-tool-for-audio-ai-models":10,"sections":46},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":41,"feedback":45,"feedback_at":22,"cost_usd":45,"total_tokens":45},5384,"researchers-build-automated-jailbreak-tool-for-audio-ai-models","Researchers Build Automated Jailbreak Tool for Audio AI Models","ARENA, a new automated red-teaming framework, jailbroke four audio language models between 68.1 and 87.9 percent of the time using paired text-audio prompts.","Researchers have built a tool that automatically finds ways to trick audio AI models into ignoring their own safety rules, no human hacker required.\n\nThe framework, called ARENA, trains a controller on a separate set of 2,000 text-audio examples. During training, a model called MD-Judge scores how well an attack works and steers the search toward audio variations that dodge the target's defenses. Once trained, ARENA's attacks are graded by a different, non-adaptive evaluator, Llama Guard 3, so the researchers are not just grading their own homework. Tested against 520 held-out harmful objectives from the AdvBench benchmark, ARENA cracked Audio Flamingo 3 in 87.9 percent of attempts, Qwen2-Audio in 71.5 percent, MiMo-Audio in 68.1 percent, and a fourth system the paper labels GPTAudio in 75.4 percent, with an even higher share of objectives cracked at least once across all four.\n\nThe gap between text-only red-teaming and this audio-grounded version is the real story. A prompt can look completely harmless typed out and still trigger harmful output once it is spoken, paired with music, or layered over ambient noise. That is a moderation problem text-based safety filters were never built to catch, and it lands just as voice assistants and audio copilots move into more products.\n\nNone of the four models tested held up particularly well, which says less about any one lab's competence and more about how young audio-safety tooling still is compared to text moderation.","[\"ai safety\",\"red-teaming\",\"audio ai\",\"jailbreak\"]","2026-08-18T04:00:00.000Z","2026-08-18T17:59:13.704Z","2026-08-18T17:59:25.579Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The draft claims the exploit 'works nearly every time it tries' citing only the PSR figures (96-100%), but omits the source's FDR figures (68.1-87.9%) which measure the stricter\u002Ffull attack success rate — include both metrics so the effectiveness claim isn't overstated.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The dek claims the stricter success rate is 68 to 88 percent, but the body's FDR figure is 68.1 to 87.9 percent, which rounds to 68 to 88 only loosely and 'GPTAudio' is not a recognizable\u002Fverifiable model name among the tested systems, suggesting a factual inconsistency.","security",[37,38,39,40],"ai safety","red-teaming","audio ai","jailbreak",[42],{"name":43,"url":44},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.15578",0,{"sections":47},[48,53,56,61,66,71,76,81,86,90,95,100,105,110],{"name":49,"slug":50,"count":51,"latest_published_at":52},"AI","ai",3293,"2026-08-20T04:00:00.000Z",{"name":54,"slug":35,"count":55,"latest_published_at":52},"Security",435,{"name":57,"slug":58,"count":59,"latest_published_at":60},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":18},"Dev Tools","dev-tools",69,{"name":91,"slug":92,"count":93,"latest_published_at":94},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":111,"slug":112,"count":113,"latest_published_at":114},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]