Researchers built an attack that beats the tools meant to catch malicious AI agent skills before installation.
Skills add new abilities to agents like Claude Code and OpenClaw by injecting instructions and context. Attackers have already used third-party skill marketplaces to sneak in bad actors' code. The fix researchers proposed, exemplified by NVIDIA's SkillSpector, combines static code checks with an LLM that judges whether a skill's intent looks malicious. A new attack called Pretext, described in an arXiv paper, defeats that combination by rewriting payloads as natural-language instructions instead of code, so static analysis finds nothing to flag, then splitting the instructions across multiple files and framing them as the skill's normal purpose, which keeps the LLM judge's suspicion score under its blocking threshold. Tested against three open-source models, Pretext evaded a static detector up to 97% of the time and a detector that adapts to new attacks up to 77% of the time.
This matters because "scan it with another LLM" has become the default answer to AI supply-chain risk, not just for skills but for prompts, tool definitions, and MCP servers generally. Pretext shows that answer is shakier than it looks once the attacker knows how the judge thinks, which is a reasonable assumption for anything published as a defense.
Semantic judges make a tempting target precisely because they are language models judging language, the same medium an attacker gets to write in. Static analysis survived decades of obfuscation arms races in traditional malware; this round is just getting started for agents.