AI/ content-moderation · vision-language-models · bluesky · ai-benchmarks

AI Models Beat Bluesky Moderation in New Benchmark Test

A new benchmark finds foundation models nearly triple the accuracy of Bluesky's own content moderation system on real posts from the platform.

Foundation models can outperform Bluesky's actual moderation system, according to a new academic benchmark - and not by a small margin.

Researchers built ModerationBench, a set of 4,000 manually annotated real-world posts pulled from Bluesky, then tested two ways of steering vision-language models to enforce content policy: one where the model reasons directly from written policy rules, and one where it learns from labeled examples of past moderation decisions. On the benchmark's random sample of posts, both approaches scored similarly at their best, and both crushed the platform's live system - an F1 score of 0.60 versus Bluesky's 0.22.

That gap matters because it is not a lab-only curiosity. Bluesky is a real platform moderating real posts right now, and the comparison is against its actual deployed system, not a strawman. If a general-purpose model can nearly triple the accuracy of a production moderation pipeline, it suggests the bottleneck for platforms isn't the raw capability of available AI - it's how that capability gets wired into policy.

The finding that instructions and examples work about equally well is the more interesting wrinkle: it hints platforms could adapt moderation to fast-changing policy without retraining on fresh labeled data every time the rules shift. Whether that holds up outside a curated benchmark, on the messier edge cases that make moderation hard in the first place, is the question this paper doesn't answer.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →