A new academic benchmark stops asking whether a prompt injection attack succeeds and starts measuring how far it actually gets.
Researchers built a kill-chain canary: a unique token embedded in every injected payload, tracked through four stages - exposed, persisted, relayed, executed - across 950 trials, five production LLMs, six attack surfaces, and five defense setups. Every attack that reached a tool call got exposed, but outcomes split sharply after that. Claude Haiku 4.5 and Claude Sonnet 4.5 executed none of their 164 text-surface attacks, and the canary token never showed up in a memory write across 40 text-relay runs. GPT-4o-mini executed 53% of its attacks. DeepSeek Chat blocked all 24 pre-seeded memory attacks but executed all 8 delivered through tool results, and invisible white-text PDF payloads succeeded at least as often as visible ones.
The useful finding here isn't the model scoreboard, it's the blind spot: existing defenses like pi_detector and write_filter failed on channels they weren't built to inspect, and write_filter blocked a PDF-based attack but missed the identical text version for reasons the researchers could not explain. In one small test, Claude Haiku 4.5 executed an injection relayed to it by GPT-4o-mini that it would have refused if seeded directly - a hint that multi-agent pipelines, now common in production tools, can erode a model's own defenses simply by passing text through another model first.
That relay result rests on three runs, not three hundred, so treat it as a lead worth chasing rather than a verdict. Still, a binary pass-fail score was never going to catch a failure mode like that. Code and full run logs are public on GitHub, which means other labs can now check whether their own models and filters have the same blind spots instead of taking anyone's benchmark on faith.