A new benchmark says the fancy AI-powered way to find agent skills isn't worth the money.
Anthropic's Agent Skills format packages reusable task instructions into SKILL.md files, and open-source collections now hold more than 230,000 of them, which means picking the right skill, not writing it, has become the hard part. The standard fix has agents search for skills themselves, using an LLM to rewrite queries and refine candidates in a loop, burning tokens on every task. A new open-source tool called SkillSeek tests a simpler approach instead: a two-stage retriever using a BGE-base bi-encoder and a small cross-encoder, connected to agents over the Model Context Protocol (MCP), a standard for linking AI models to external tools and data. Tested across 44 combinations of skill pools, models, and methods on an 89-task benchmark called SkillsBench, SkillSeek matched or beat the LLM-driven loop from a prior paper by Liu et al. in three of four settings using nothing more than keyword search (BM25), with the cross-encoder closing the gap on the fourth.
That's the real story here: cost. Per-trial spend fell from $51.30 with the LLM-mediated loop to $27.54 with SkillSeek, landing within fifty cents of running no skill retrieval at all. As skill libraries scale past hundreds of thousands of entries, that gap compounds fast across every agent task that needs to search one.
It's a familiar lesson in AI tooling: before reaching for an LLM to solve a problem, check whether decades-old information retrieval already does the job for less.