Meta says it has fixed a security fix that was quietly breaking AI agents.
Researchers at Meta and academic collaborators have released Meta-SecAlign, a new fine-tuning method for defending large language models against prompt injection - the attack where hidden instructions in a webpage, email, or document hijack an AI agent's behavior. It's built to fix a problem with SecAlign, an existing open-source defense: at larger scale, the team found SecAlign quietly tanked model usefulness, especially on agentic tasks like tool-calling and web navigation. Meta-SecAlign instead trains on injected prompts placed at randomized positions, so the model can't just learn a shortcut, and uses the model's own generated answers as training labels rather than hand-written ones. The team tested it on five models - including Meta's Llama 3.1, 3.3, and 4 Scout, plus two from Alibaba's Qwen3 family - across six benchmarks including AgentDojo, InjecAgent, WASP, and SEP. One listed model, described as "Qwen3.6-27B," doesn't match any released Qwen3 model - Alibaba's published lineup tops out at 4B, 8B, 14B, and 32B versions - and the paper offers no explanation, which looks like an error or an unannounced variant.
Prompt injection is the vulnerability security researchers keep calling the top open problem for AI agents, since it lets attackers turn an assistant's own tool access against its user. Most defenses trade security for capability, making agents safer but too neutered to be useful - exactly the tradeoff Meta says it avoided here. If the utility numbers hold up outside Meta's own testing, that matters more than the security scores alone, since it removes the usual excuse for skipping defenses in production agents.
Meta published the code and two model checkpoints on Hugging Face for anyone to check the claims - though an unexplained model name in the paper's own table is a good reason to check them closely.