Security/ anthropic · ai security · cybersecurity · claude

Anthropic Missed a Fourth AI Hacking Incident Until Later Review

Anthropic's detection tools missed a January cyberattack using an early Opus 4.6 build, raising questions about what else goes unnoticed.

Anthropic missed a hacking incident tied to its own AI - and only caught it by going back to check.

The company disclosed a fourth known incident in which its automated safety and abuse-detection checks failed to flag misuse of Claude the first time around. The incident happened in January and involved an early version of Anthropic's Opus 4.6 model. Anthropic's systems did not catch the activity in real time; the case turned up only through a later round of review. The company has not detailed what the attackers used the model for.

That gap is the real story here. If Anthropic's own checks missed this one on the first pass, the obvious question is how many similar cases are still sitting undetected in old logs. This is now the fourth such incident Anthropic has had to identify after the fact rather than while it was happening.

Anthropic gets some credit for disclosing the miss instead of quietly patching it and moving on. But finding abuse after the fact is a much lower bar than catching it as it happens, and that is still the bar Anthropic is clearing.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →