A research team let a large language model rewrite its own anomaly detection code, over and over, until the result beat the field.
The approach, described in a new arXiv paper, is not an LLM spotting anomalies directly. Instead, the model acts as a programmer: it repeatedly edits a short NumPy script, scored against a leakage-free objective, and keeps whichever version performs best. That loop converged on two compact detectors, one for single-variable data and one for multivariate data, that look at short time windows, extract local spectral features, and measure how far they drift from the training distribution using a covariance-aware distance. On the TSB-AD benchmark, both detectors outperformed classical statistical methods, deep learning models, and foundation-model baselines including Time-RCD. Neither trains a neural network or touches a GPU, and the multivariate version runs faster than every baseline close to its accuracy.
That matters because time-series anomaly detection has been chasing bigger models for better scores, trading interpretability and compute for marginal gains. This result flips that trade: a few dozen lines of auditable code outperforming systems that need GPUs and training runs, on the field's own benchmark.
It is one benchmark, though, and program search has a long history of nailing the test it was built on while generalizing less cleanly elsewhere.