A new open-source pipeline called SkillFlow helps AI coding agents pick the right pre-written skill out of a 35,000-entry library instead of choking on irrelevant ones.
SkillFlow indexes roughly 35,000 community-contributed SKILL.md files scraped from GitHub, then narrows that pool through four stages: dense retrieval, two rounds of cross-encoder reranking, and a final LLM-based selection pass. Researchers tested it on two coding benchmarks. On SkillsBench, giving agents SkillFlow-retrieved skills raised their Pass@1 rate from 9.2% to 16.4%, a 78.3% relative increase that closed most of the gap to an 84.1%-of-oracle ceiling. On Terminal-Bench, agents happily used the retrieved skills 70.1% of the time, but their scores did not budge.
That split result is the real finding here. Good retrieval only helps if the library actually contains runnable, well-built skills for the task at hand - on Terminal-Bench, it apparently did not. As agent frameworks race to build plugin marketplaces and skill stores, this is an early data point that indexing more skills matters less than indexing better ones.
Also worth sitting with: even the improved number means the agent still fails to solve a task correctly more than four times out of five.