A new academic defense claims it can stop large language models from obeying instructions smuggled inside the data they are asked to read, without retraining the model at all.
Researchers describe Learnable Trust-Boundary Delimiters (LTBD), a lightweight method that wraps trusted user instructions and untrusted external content in a small set of learnable markers, teaching the model which parts of its input it should actually listen to. The underlying model's weights are left untouched. On the AlpacaFarm benchmark, the method reported a 0.00% attack success rate; on TaskTracker, it reported between 0.11% and 0.19%. Those numbers held even in tests where the simulated attacker knew exactly how the defense worked and tried to design around it.
Prompt injection is the practical problem lurking behind splashy AI agent demos: a model told to summarize a webpage or read an email can get hijacked by instructions planted inside that content, because today's LLMs don't reliably separate 'my user told me to do this' from 'some webpage says to do this.' Most existing fixes force a tradeoff, either costly fine-tuning that works until an attacker finds a new blind spot, or handcrafted prompt wrappers that crumble after a few rounds of probing. A cheap, parameter-free layer that still holds up against an attacker who already knows the trick would be a genuinely useful floor to build on.
Benchmark numbers from a paper's own authors testing their own defense are not the same as what happens once outside red teams get their hands on it, and prompt injection defenses have a long history of looking solid right up until they didn't.