A new testing framework just caught dozens of bugs hiding in the instructions that run Claude Code, Codex CLI, and Gemini CLI.
Researchers built a tool called Arbiter that treats system prompts, the hidden instructions that govern an AI coding agent's behavior, as software worth testing instead of prose worth trusting. It pairs formal evaluation rules with several different LLMs cross-checking the same prompt for contradictions. Run against the prompts behind Anthropic's Claude Code, OpenAI's Codex CLI, and Google's Gemini CLI, it surfaced 152 findings in an open-ended scan, then 21 hand-labeled interference patterns, cases where one instruction quietly undercuts another, in a deeper pass on a single vendor. The entire three-way analysis cost $0.27.
A prompt's architecture, whether it is one dense block, a flat list, or broken into modules, predicted the kind of failure that turned up, though not how severe it was. More notable: checking a prompt with multiple models surfaced entirely different bug categories than checking it with one, which suggests single-model prompt reviews are missing whole classes of problems.
One finding was a data-loss bug in Gemini CLI's memory system that matches an issue Google already patched. The patch fixed the symptom, not the schema-level flaw Arbiter identified, which is a decent summary of where AI tooling fixes stand right now: triage over root cause.