AI/ ai-agents · llm-tooling · agent-harness · arxiv

STITCH Builds Custom AI Agent Harnesses on the Fly

A new framework swaps fixed agent setups for mixed-and-matched reusable parts, lifting task success up to 12 points over static baselines.

Researchers have built a system that assembles a custom toolkit for an AI agent right before it tackles a task, instead of forcing every job through the same fixed setup.

The approach, called STITCH, targets "harnesses" - the scaffolding that controls how an AI model gathers context, calls tools, checks its own work, and knows when to stop. Today most agents run on one harness for every task, even though a mechanism that helps on one job can slow down or distract the model on another. The researchers mined failed task attempts to extract "Harness Primitives," reusable mechanisms with defined scopes, then built STITCH to pick and assemble the right primitives for each task at test time, skipping the cost of writing or debugging new harness code on the fly.

This matters because agent performance gains have increasingly come from scaffolding, not just bigger models, and a one-size-fits-all harness is a bottleneck nobody talks about. STITCH reportedly beat fixed baselines by up to 12 points on task success and outperformed Codex CLI, a human-designed harness, while adding only 2.7 percent overhead - a fraction of the cost of generating a bespoke harness from scratch.

If the numbers hold up outside the paper's own benchmarks, this looks less like a clever trick and more like an admission that "one harness to rule them all" was always going to hit a ceiling.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →