AI/ ai safety · open-weight models · uncensored ai · huggingface

Uncensored AI Models Keep Spreading After Takedowns

A new study of 3,471 uncensored AI models found quantized copies persist across mirrors after takedowns, and a quarter of downstream apps are malicious.

Stripping safety filters off open-weight AI models has become a cottage industry, and deleting the original doesn't make it go away.

A new arXiv study tracked the uncensored open-weight model ecosystem on HuggingFace from January 2024 to March 2026. Researchers identified 3,471 original uncensored models, each repackaged an average of 2.4 times, yielding 8,164 compressed redistributions. Three actors account for 52% of those redistributions. Once a model is quantized and mirrored across different accounts, file formats, and registries such as Ollama, it keeps circulating even after the original upload is taken down.

The persistence problem gets worse downstream. The researchers found 1,643 GitHub projects built on uncensored models, and classified 25% of them as explicitly malicious. That means moderation aimed at the source models misses most of the actual risk, which lives in the derivative copies and the applications built on top of them.

It's the same lesson piracy enforcement learned decades ago: once something is copied and small enough to run locally, takedowns are more theater than defense.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →