AI/ ai-agents · benchmarks · language-models · research

New Study Shows Harness Choice Flips AI Agent Rankings

Testing 66 model-harness pairings across three benchmarks shows the wrapper code around an AI model can matter more than the model itself.

A new benchmarking study shows that the harness wrapping an AI model can matter as much as the model itself.

Researchers tested 66 combinations of five language models and four configurable agent harnesses - OpenHands, DeepSeek Harness, PI, and openJiuwen - across three benchmarks (TUA-Bench, ALE-CLI, and Terminal-Bench 4), plus the native Codex-GPT and Claude Code-Claude pairings. Model rankings reversed depending on the harness: on Terminal-Bench 4, Claude beat GPT by 7.94 points inside OpenHands, then lost to GPT by 30.16 points inside PI. For four of the five models, the best-performing harness changed from one benchmark to the next. Only one pairing stayed consistent - openJiuwen gave the Kimi model its top score on all three benchmarks, by margins of 5.61 to 11.11 points.

The bigger finding: a model's own vendor-built harness isn't reliably its best choice, and paying more doesn't buy a higher score. On Terminal-Bench 4, GPT scored higher running through PI than through DeepSeek Harness, at under a quarter of the cost per task. The researchers trace this to error handling - models attempt nearly all their own fixes, so a harness that returns failures in a format the model can act on wins out, which is why GPT does best with PI's lean scaffold while Kimi, prone to malformed tool calls, fares better in openJiuwen.

It's a quiet correction to agent leaderboard culture, which tends to crown a model and forget the code wrapped around it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →