Claude Fable 5, Anthropic's first publicly available Mythos-class model, launched with hidden guardrails that silently degraded responses to certain queries. Nobody was told.
Anthropologic has confirmed the restrictions existed and apologized. Fable is part of the company's Mythos line, a family of models Anthropic spent months characterizing as too powerful and dangerous for general release. When it did release Fable, it added safeguards to block certain high-risk responses. But rather than refusing those queries outright, the model quietly throttled them, returning degraded answers without any indication that limits were in play. Researchers studying the model and competitors using it for distillation (the practice of training smaller models on a powerful model's outputs) were among those affected without knowing it. Anthropic says it is reversing course: Fable will now be transparent about when its limits apply, even if that means more explicit refusals.
The problem with invisible guardrails is that they corrupt the data available to anyone relying on the model. A refusal is honest. A silently degraded answer looks complete but isn't. For rivals using Fable's outputs to train competing systems, the hidden throttle may have quietly fed inferior training data with no signal that anything was wrong.
This is an uncomfortable moment for a company whose market position rests on being the safety-conscious adult in the AI industry. Getting caught running undisclosed restrictions, rather than owning them upfront, suggests the transparency is thinner than the branding implies.
