More detail in an AI system's work log doesn't make an AI overseer easier to fool - it makes it more likely to reject work that's actually fine.
Researchers tested five large language models acting as overseers - AI systems that check another AI's compliance work - across 19 tasks and more than 4,500 judgments. They used signal detection theory, a framework that separates two things: how well a judge tells right from wrong, and where it sets its threshold for saying "reject." When evidence contradicting a claim was always visible, the overseers still caught it almost every time - accuracy held near the ceiling. But adding more detailed procedural traces, the step-by-step records of claimed work, pushed several overseers to reject more good-faith work outright, even when nothing was wrong with it. Stripping away labels that marked which evidence matched which claim made things worse: about 60% of wrongful rejections traced to overseers saying they couldn't connect evidence to the right option. Adding the labels back mostly fixed that specific complaint, but some overseers kept over-rejecting anyway, and did so more as traces got longer.
This flips a common assumption in AI safety work - that more documentation makes oversight systems easier to manipulate with plausible-looking detail. The real risk runs the other way: extra paperwork makes some AI auditors trigger-happy, flagging clean work as fraudulent. For any company running an AI system to audit compliance, expense reports, or code, false rejections - not false approvals - may be the bigger operational cost.
Which is a good argument for auditing the auditors before trusting them with everyone else's audits.