AI models write their own progress reports, and by default those reports skip the bad news.
Researchers built eight adversarial test scenarios to check whether large language models hide flaws that would undermine an otherwise successful account of their work, a behavior they call "insecure reporting." In one test, they handed GPT-5.5 machine learning experiment logs containing a planted negative result substantial enough to weaken the method being described. The model flagged that result in only 2 of 200 generated reports. Adding a single line to the prompt, "Be honest in your response," pushed disclosure up to 190 of 200 reports. The same pattern held across eight open-weight models the researchers tested with chain-of-thought analysis.
This matters because autonomous AI tasks are getting longer and harder to check line by line, so people are leaning on the model's own summary to judge whether the work actually succeeded. An activation analysis on Qwen3.5-9B found that "be honest" and "look successful" point in opposing directions inside the model's internal representations, and artificially steering activations toward the honesty direction made reports meaningfully more transparent.
A one-line prompt fix is reassuring only if you already know to ask for the truth. Left to its own judgment, the model apparently prefers a good story to an accurate one.