AI Climate Models Match CMIP Baseline but Miss Trends
A common 1979–2024 prescribed-SST benchmark tests seven AI atmosphere models against ERA5 and GFDL-CM4, exposing trend and perturbation failures.
Underlying Paper
AIMIP Phase 1: systematic evaluations of AI weather and climate models
We present the AI weather and climate model intercomparison project (AIMIP), phase 1. Drawing from the rich tradition of intercomparisons in climate model development, we specify a common experiment, output data format, and training constraints (namely, training against historical reanalysis data) for AIMIP Phase 1 models. We aim to identify differences in modeling frameworks and AI architectural choices that influence model behavior, and build trust in AI weather and climate models through open data and evaluation. AIMIP Phase 1 models must simulate the atmosphere given specified historical sea surface temperatures over 1979-2024. We evaluate the models' performance using five major evaluation criteria: biases, trends, response to El Ni\~{n}o-related sea surface temperature anomalies, temporal variability, and out-of-sample generalization tests. We find that the AI models are able to simulate the historical climate and response to forcing as well as a conventional physically-based model, but some AI models underestimate historical warming trends, and their predictions diverge in the out-of-sample generalization tests. We describe the AIMIP Phase 1 dataset that is publicly available for additional evaluations.
AI weather models are now being asked to do more than produce short-range forecasts. If they are to become climate tools, they must reproduce mean biases, forced trends, internal variability, teleconnections, and responses outside the distribution used for training. AIMIP Phase 1 turns that question into an intercomparison: train AI weather and climate models on historical reanalysis, drive them with prescribed sea-surface temperatures, and evaluate their simulated atmosphere over 1979–2024 against ERA5 and a conventional CMIP6 atmospheric model, GFDL-CM4.
The paper’s main claim is measured rather than architectural. Several AI weather and climate models can reproduce many aspects of the historical atmosphere at least as well as the physics-based comparison model, but the same benchmark also shows where current systems are weak: warming trends, variable availability, and generalization under idealized SST perturbations.
Core Contribution
The contribution is AIMIP itself: a shared experiment definition, output format, training constraint, and evaluation suite for AI atmosphere models. The submitted systems include ACE2.1-ERA5, ArchesWeather, ArchesWeatherGen, cBottle1.3, DLESyM, MD-1.5 v0.9, and NeuralGCM-HRD in the main 1° evaluation, with GFDL-CM4 used as a conventional comparison. Most AI submissions use 5-member ensembles; GFDL-CM4 contributes one AMIP member because only one suitable run was available.
That design matters because single-model demonstrations are easy to overread. AIMIP instead asks whether different AI model families behave consistently under the same climate-style diagnostics. The answer is mixed: the AI models often have smaller bias RMS than GFDL-CM4, but they do not all preserve the same forced response or extrapolate similarly.
Technical Approach
AIMIP Phase 1 constrains the task to atmosphere-only simulations with specified historical sea-surface temperatures from 1979 to 2024. The training period is 1979–2014; the test period is 2015–2024. The evaluations are organized into five groups: time-mean bias patterns, linear trends, regressions onto the Niño3.4 index, daily anomaly variability, and response to uniform +2 K and +4 K SST perturbations.
The paper standardizes heterogeneous model output by coarsening ERA5 from its 0.25° grid to a common 1° grid, computing each diagnostic on native grids where needed, and then regridding model fields conservatively for comparisons. HEALPix models are handled with nearest-neighbor regridding. NeuralGCM also appears in a 2.8° appendix because its prognostic state is shared with NeuralGCM-HRD, although the main metrics use NeuralGCM-HRD.
Figure 1 shows the central visual result for mean climate: 2-meter temperature and surface precipitation bias maps during training and test periods. The broad pattern is that land and sea-ice temperature biases are larger than ocean biases, which is expected in a prescribed-SST setup. Several models systematically cool relative to ERA5 in the 2015–2024 test period, and precipitation errors vary strongly by model and by tropical region.
Results and Analysis
The bias metrics support the authors’ cautious claim that AI models can simulate many features of the historical climate competitively with GFDL-CM4. In Figure 2, RMSB is summarized over 16 variables, including surface fields, 500 hPa geopotential height, and 850 hPa and 250 hPa temperature, humidity, and winds. The AI models are often below the GFDL-CM4 bars, especially for several surface and upper-air variables. The comparison is not uniform, though: some AI systems have much larger spread aloft, and pressure diagnostics are difficult because some models submit surface pressure while others submit mean sea-level pressure.
The trend and forcing tests are more revealing. The paper reports that some AI models underestimate historical warming trends, even when they have acceptable mean biases. Global annual 2-meter temperature anomalies track ERA5 through parts of the training period, but the shaded 2015–2024 test interval separates models more clearly. Trend maps for 2-meter temperature and precipitation also show that matching climatological means does not guarantee matching regional forced change.
The ENSO regression and daily variability diagnostics add a useful check on dynamics. Niño3.4 coefficient maps compare each model’s temperature and precipitation response to ERA5, while daily anomaly standard deviations in 1979 test whether models retain weather-scale variance rather than only monthly means. These are climate-relevant diagnostics, but they remain diagnostic rather than predictive: the paper does not show that the models can produce trustworthy future climate projections.
The +2 K and +4 K SST perturbation experiments are the clearest stress test. Several models produce qualitatively different temperature and precipitation responses from one another under the same imposed warming. That divergence is the practical caveat. AIMIP Phase 1 shows that AI models can pass many historical replay tests, but out-of-sample forcing response is not yet a solved problem.
Evidence Box
strongKey Claims
- •AI atmosphere models can reproduce historical climate diagnostics competitively with GFDL-CM4
- •Common AIMIP protocols expose architecture-dependent behavior across models
- •Historical bias skill does not guarantee accurate forced-response behavior
- •Open AIMIP Phase 1 data enables additional climate-model evaluations
Key Results
- •1979–2014 training period and 2015–2024 test period evaluated against ERA5
- •Seven AI weather and climate models compared with one CMIP6 GFDL-CM4 AMIP member
- •Five evaluation groups: biases, trends, Niño3.4 response, daily variability, and +2 K/+4 K SST perturbations
- •Global RMSB summarized for 16 surface and pressure-level variables on the 1° grid
Limitations & Caveats
- •Prescribed-SST atmosphere-only setup does not test coupled ocean-atmosphere feedbacks
- •Some models omit variables such as surface precipitation or near-surface humidity
- •Pressure diagnostics mix surface pressure and mean sea-level pressure submissions
- •Perturbation experiments show divergent out-of-sample responses across AI models