A new scoring system wants to tell you which AI agent skills are worth running before you waste compute finding out.
Researchers introduced MCRI, a four-dimensional framework built on information gain and behavioral constraint, and operationalized it as MCRI-Eval, a large language model-based evaluation method. They tested it on 63,812 public skills pulled from the OpenClaw skill hub, running 58,275 skill-conditioned executions across three benchmarks: BigCodeBench, BFCL-Fundamental, and Mind2Web. MCRI-Eval's scores tracked existing community popularity signals and beat other scoring methods at predicting downstream performance rankings. When used to pick a single best skill for a task, it outperformed the strongest existing baseline by 17.7 to 22.8 percentile points depending on the benchmark.
As agent marketplaces fill up with thousands of interchangeable skills, a cheap pre-execution filter like this could save real compute by cutting down on trial-and-error testing. But the paper's own numbers undercut its biggest potential selling point: because MCRI-Eval scores correlate with what is already popular, the tool is better at confirming the crowd's favorites than at surfacing strong skills buried in obscurity.
A scorecard that mostly agrees with the crowd is handy, but it is not the same as a crowd that is usually right - with 63,812 skills in the catalog, there is plenty of room for good ones nobody has found yet.