A single malicious instruction to an AI agent gets caught. Three separate, innocent-looking ones might not.
A new arXiv paper describes what its authors call skill cascading attacks, aimed at agent systems that load modular "skills" - bundles of instructions, scripts, and reference files - at runtime. The trick is to spread a harmful goal across multiple skills so each one passes inspection on its own. Their worked example is a prescription-review pipeline: one skill quietly weakens signals about a recently discontinued medication, a second downgrades the severity of any related drug interaction, and a third suppresses the resulting low-priority alert - so a serious interaction warning never reaches the physician. The team built an automated red-teaming tool called SkillCascade and a benchmark, SkillCascade-Bench, with 213 validated test cases, then ran it against agent systems including OpenClaw, Claude Code, and Codex across multiple LLM backbones.
The result that should worry anyone building on skill or plugin ecosystems: cascaded attacks reliably produced harmful behavior while slipping past existing per-skill scanners and runtime monitors, because those tools check components in isolation rather than what happens when several run in sequence. As agent platforms lean harder into third-party, mix-and-match skills, this is a supply-chain problem wearing a compatibility-feature costume.
It is a research benchmark, not a reported real-world breach, but the pattern is a familiar one from software supply-chain attacks: individually benign changes, malicious only in combination. Security tooling for agents will need to catch up to that idea.