Making large language models answer faster also makes them easier to jailbreak.
Researchers ran the first systematic security study of speculative decoding, the technique that pairs a small "draft" model with a larger "target" model to speed up inference. Across a wide range of lossy speculative decoding methods, they found a sharp asymmetry: the more a method sped things up, the more jailbreak and prompt injection attacks succeeded against it. Attack success rates climbed much faster than the models' actual usefulness declined. The team traced the weakness to early tokens generated by the draft model, which tend to get accepted without enough scrutiny.
That matters because speculative decoding is already a routine trick for cutting inference costs, which means production systems may be quietly trading safety for speed without anyone measuring the tradeoff. The researchers' proposed fix, called SecureSD, applies a stricter verification check specifically to those early draft-model tokens, improving security while preserving most of the efficiency and utility gains.
It's a useful reminder that "faster" and "safe" are not the same axis, and that optimizing one without auditing the other is exactly how security holes ship quietly into production.