A researcher just showed you can jailbreak six popular open-weight language models by throwing random noise at their prompts.
The attack, called Perturbed Embedding Vector, skips the usual jailbreak playbook entirely. Instead of crafting clever prompts or running gradient-based optimization against a model's weights, it just adds random Gaussian noise directly to the embedding vectors that represent a prompt, then resamples until something breaks through. Tested against six open-weight LLMs on the JailbreakBench benchmark, it produced an unsafe response for every single prompt on every model. The first successful break typically landed within a minute, using up to ten times less compute than prior attack methods.
That speed and simplicity matter more than the headline number, because safety fine-tuning is supposed to make a model refuse harmful requests no matter how the prompt is phrased. PEV never touches the phrasing; it perturbs the internal representation instead, and the refusals collapse anyway. For open-weight models, where anyone can load the weights and poke at embeddings directly, that's a much easier attack surface to reach than the proprietary APIs most prior jailbreak research targeted.
Open weights were supposed to be the transparent, inspectable alternative to closed systems; this paper is a reminder that transparency cuts both ways.