A new study tested anomaly detection systems against real data from a nuclear power plant, and the results complicate the usual pitch for fancier AI.
Researchers compared streaming anomaly detection methods with state-of-the-art time series anomaly detection (TSAD) models, running both online against a real dataset from a nuclear plant operated by EDF, the French utility that depends on continuous monitoring to catch problems as they happen. The team also tested automated anomaly detection tools in that same live setting, rather than the offline benchmarks most such papers rely on. No single method dominated across the board. But TSAD models produced more consistent results when run online, and combining multiple models through ensembling added robustness that no single detector matched on its own.
That distinction matters because most anomaly detection research is tested on static, cleaned-up data, not a live industrial stream where new readings keep arriving and there is no going back to relabel the past. A nuclear plant is about as high-stakes a proving ground as this kind of monitoring gets, which makes the online results more credible than another synthetic benchmark.
The takeaway: in a safety-critical pipeline, a boring ensemble of detectors beat any single flashy model.