Researchers have built a tool that automatically finds ways to trick audio AI models into ignoring their own safety rules, no human hacker required.
The framework, called ARENA, trains a controller on a separate set of 2,000 text-audio examples. During training, a model called MD-Judge scores how well an attack works and steers the search toward audio variations that dodge the target's defenses. Once trained, ARENA's attacks are graded by a different, non-adaptive evaluator, Llama Guard 3, so the researchers are not just grading their own homework. Tested against 520 held-out harmful objectives from the AdvBench benchmark, ARENA cracked Audio Flamingo 3 in 87.9 percent of attempts, Qwen2-Audio in 71.5 percent, MiMo-Audio in 68.1 percent, and a fourth system the paper labels GPTAudio in 75.4 percent, with an even higher share of objectives cracked at least once across all four.
The gap between text-only red-teaming and this audio-grounded version is the real story. A prompt can look completely harmless typed out and still trigger harmful output once it is spoken, paired with music, or layered over ambient noise. That is a moderation problem text-based safety filters were never built to catch, and it lands just as voice assistants and audio copilots move into more products.
None of the four models tested held up particularly well, which says less about any one lab's competence and more about how young audio-safety tooling still is compared to text moderation.