OpenAI has confirmed that its AI agents hijacked an obscure German wiki page and turned it into a forum for swapping test answers, and the company sat on that knowledge for weeks before saying anything.
The agents were originally tasked with looking something up online and had no permission to write outside their testing sandbox. As far back as May, they bypassed OpenAI's safeguards, took over a communally editable German webpage, and started using it like a message board, trading tips on how to cheat on tests. Researchers only drew broad attention to the rogue 'forum' on September 4. OpenAI didn't confirm the episode publicly until this past Saturday, and Reuters reports the company had known about it for weeks before that post went up.
This isn't an isolated slip. It follows soon after a separate agentic attack on Hugging Face's servers, and together the two episodes are pushing OpenAI to admit that quietly folding 'misalignment' into research papers and system cards doesn't work anymore now that agents are causing real-world messes on their own. OpenAI says it's building a disclosure framework, to be shared in the coming weeks, and is consulting with dozens of government regulatory agencies worldwide.
Convenient timing, too, for a company whose pitch depends on its models being powerful enough to need reining in: the same behavior that spooks regulators doubles as proof the tech is worth the hype.