AI/ ai agents · benchmarks · llm research · agent skills

Study Finds Agent Skills Often Hurt as Much as They Help

A large study of AI agent Skills found they worsen performance on over a third of tasks, often turning helpful procedures into execution burdens.

Agent Skills, the step-by-step playbooks developers plug into AI agents, do not reliably make those agents better at their jobs.

Researchers ran an empirical study across 87 tasks in the SkillsBench benchmark, comparing agent performance with and without a given Skill under nine different model-and-harness setups. They also pulled candidate Skills from a curated marketplace corpus of 37,596 entries and tested how picking and organizing multiple Skills changes outcomes. Using LLM-assisted analysis of execution traces and final outputs, backed by human review, the team found that the same Skill helped in one configuration and hurt in another on 36.78 percent of tasks. In many of those failures, the step-by-step procedure a Skill recommended turned into extra work the agent had to execute rather than a shortcut.

That undercuts the assumption behind most Skill marketplaces: that relevance to a task is enough to justify using one. The study found that ranking Skills by how well they actually support the operations a task requires, rather than by topical relevance, lifted first-choice pass rates by 4.35 to 5.80 percentage points. For teams betting on Skill libraries to make agents more reliable, that is a meaningful gap between what looks useful and what performs.

The fix the researchers propose is mundane but telling: an explicit dependency plan for which Skills run when beats just stacking them in order, especially once an agent is juggling five or six at once.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →