A subtle quirk in how time series forecasting models are trained has been quietly skewing their accuracy, and the fix is a single line of code.
Time series foundation models learn from huge, mixed collections of data with wildly different scales, think server response times measured in milliseconds sitting next to national GDP figures in the trillions. A common technique called Reversible Instance Normalization rescales each series before training, then reverses that scaling before the loss is calculated. Researchers show that reversing the transform multiplies each series' gradient by a factor tied to its own scale, so series with bigger numbers end up dominating training by default. They label this failure mode scale-contaminated training, and prove that computing the loss on the already-scaled values instead makes the entire training process mathematically invariant to how any individual series happens to be scaled.
The fix is not marginal. Tested across four different model architectures on standard forecasting benchmarks, it cut forecasting error by an average of 18.8% on one benchmark suite and 21.9% on another, and it also improved accuracy in most of the supervised forecasting setups tested. That is a significant accuracy gain coming purely from how a loss function is computed, not from more data or a bigger model.
Worth noting that existing models do not even agree on which version of this calculation they use, and rarely report it, which suggests this bug has been sitting unflagged in benchmark results for a while.