AI models that forecast weather and climate are being judged by a test that might not measure what matters.
A paper posted this week on arXiv examines foundation models trained on weather and climate data that are now being fine-tuned for Earth science tasks well beyond forecasting. The authors say development of these models is outpacing the field's ability to evaluate them. Right now, models are scored almost entirely on benchmark skill: how closely a forecast matches a reference dataset, not whether it understands the physical processes it is simulating. The paper lays out five priorities for physical evaluation, covering training data, fine-tuning, behavioral testing, mechanistic interpretability, and output validation, and recommends three steps for the next decade: open AI-ready evaluation datasets, a shared standard for physics-based evaluation, and a dedicated research program on model safety.
That gap matters because climate change keeps shifting conditions away from anything these models were trained on. A model can score well by reproducing patterns from the past and still misrepresent the physics that will govern a hotter, less stable climate. Treating benchmark accuracy as a stand-in for physical correctness bets that past performance predicts future reliability, in exactly the conditions where that assumption breaks down.
It is a familiar problem in new clothes: pattern-matching getting mistaken for understanding, the same critique long leveled at large language models. Earth science just has higher stakes when the model is wrong.