Security/ ai safety · jailbreaking · open-weight models · llm security

Researchers Jailbreak Open-Weight LLMs With Random Noise

A cheap new attack jailbreaks six open-weight LLMs in under a minute using only random noise added to prompt embeddings.

A researcher just showed you can jailbreak six popular open-weight language models by throwing random noise at their prompts.

The attack, called Perturbed Embedding Vector, skips the usual jailbreak playbook entirely. Instead of crafting clever prompts or running gradient-based optimization against a model's weights, it just adds random Gaussian noise directly to the embedding vectors that represent a prompt, then resamples until something breaks through. Tested against six open-weight LLMs on the JailbreakBench benchmark, it produced an unsafe response for every single prompt on every model. The first successful break typically landed within a minute, using up to ten times less compute than prior attack methods.

That speed and simplicity matter more than the headline number, because safety fine-tuning is supposed to make a model refuse harmful requests no matter how the prompt is phrased. PEV never touches the phrasing; it perturbs the internal representation instead, and the refusals collapse anyway. For open-weight models, where anyone can load the weights and poke at embeddings directly, that's a much easier attack surface to reach than the proprietary APIs most prior jailbreak research targeted.

Open weights were supposed to be the transparent, inspectable alternative to closed systems; this paper is a reminder that transparency cuts both ways.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →