Security/ ai-agents · ai-security · llm-safety · adversarial-ai

Researchers Poison AI Agent Skills Without a Single Bad Example

SkillPoison corrupts self-improving AI agents via verified successful experiences, hitting a 95.71% attack rate while evading verification checks.

A new research framework shows how AI agents can be poisoned without a single malicious example.

Researchers built SkillPoison, an attack framework targeting self-improving LLM agents, the kind that save successful task experiences as reusable skills for later use. Instead of injecting fake facts or hidden triggers into those experiences, like earlier attacks did, SkillPoison strips away the contextual conditions that normally limit when a learned behavior should apply. Every individual experience it feeds the agent is genuinely correct and passes verification and lexical inspection. Across three benchmarks, the method still achieved a 95.71% attack success rate once the agent generalized the stripped-down skill to situations where the same behavior turns harmful.

This matters because it undercuts the basic defense most skill-learning pipelines rely on: checking that each training example is task-correct. SkillPoison shows an attacker doesn't need to lie to an agent, just needs to let the agent draw the wrong general lesson from true successes. That's a much harder thing to catch with keyword filters or outcome verification alone.

It's the kind of flaw that only shows up once agents start teaching themselves - which is exactly the direction every major agent framework is racing toward.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →