AI/ ai agents · benchmarking · skill evaluation · research

A New Way to Score AI Agent Skills Before Running Them

A new scoring framework predicts which AI agent skills perform best before you run them, but it still leans heavily on existing popularity data.

A new scoring system wants to tell you which AI agent skills are worth running before you waste compute finding out.

Researchers introduced MCRI, a four-dimensional framework built on information gain and behavioral constraint, and operationalized it as MCRI-Eval, a large language model-based evaluation method. They tested it on 63,812 public skills pulled from the OpenClaw skill hub, running 58,275 skill-conditioned executions across three benchmarks: BigCodeBench, BFCL-Fundamental, and Mind2Web. MCRI-Eval's scores tracked existing community popularity signals and beat other scoring methods at predicting downstream performance rankings. When used to pick a single best skill for a task, it outperformed the strongest existing baseline by 17.7 to 22.8 percentile points depending on the benchmark.

As agent marketplaces fill up with thousands of interchangeable skills, a cheap pre-execution filter like this could save real compute by cutting down on trial-and-error testing. But the paper's own numbers undercut its biggest potential selling point: because MCRI-Eval scores correlate with what is already popular, the tool is better at confirming the crowd's favorites than at surfacing strong skills buried in obscurity.

A scorecard that mostly agrees with the crowd is handy, but it is not the same as a crowd that is usually right - with 63,812 skills in the catalog, there is plenty of room for good ones nobody has found yet.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →