An arXiv paper this week argues the code wrapped around a large language model deserves as much design attention as the model itself.
The paper introduces Meta-Harness, a system that treats harness design - the prompts, call routing, and output parsing that sit between a user and an LLM - as a search problem with three competing goals: accuracy, safety, and token cost. The search is run by an agentic proposer, Claude Code, given full filesystem access to prior harness code, execution traces, and scoring data. The resulting method, MoMHa, beat ten baselines across seven synthetic test suites (a joint score of 0.482 versus 0.198 to 0.422) and outperformed the strongest existing baseline, a framework called DSPy, on seven real-world benchmarks that include HumanEval, MBPP, and MMLU-Pro (0.461 versus 0.377). It also posted the best safety score on three benchmarks built from U-SafeBench and used 95 fewer tokens per example than a two-step version of the same idea.
Most harness-tuning work still optimizes for accuracy alone and treats safety and cost as afterthoughts. This paper's contribution is making the trade-off explicit and showing that optimizing all three at once beats doing it in stages. It's also notable that the tool doing the optimizing, Claude Code, is a coding agent being pointed at the scaffolding around other models rather than at application code - a use case that goes beyond typical software tasks.
The authors say the strategies transferred to benchmarks the system never trained on for most of the twelve models tested, which is a strong transfer claim. The team plans to release the harness code and evaluation logs, so that number will get tested against reality soon enough.