WP-MIP Standardizes Weather Model Comparisons

A WMO-backed protocol pools physical, AI, and hybrid forecasts across six continents to test skill, physical consistency, and operational usefulness.

Editorial Desk·July 28, 2026·4 min readmoderate

Underlying Paper

WP-MIP: An Artificial Intelligence, Hybrid, and Physically Based Model Intercomparison Project for Weather Prediction

Rapid progress in the field of machine-learning for weather prediction has led to the emergence of algorithms whose forecasting skill can exceed that of traditional physically based models. This development represents an opportunity to improve the quality of forecasting services provided by operational centers, particularly given the speed at which machine-learning based models generate predictions. Despite the clear promise of these systems, questions remain about the ability of the current generation of machine-learning models to generate physically consistent predictions of the full suite of required forecast fields under all conditions. Answering these questions will require careful comparisons between the well-understood physically based models, current state-of-the-art machine-learning models, and the hybrid models that combine elements of these two archetypes. The Weather Prediction Model Intercomparison Project (WP-MIP) is a World Meteorological Organization-supported initiative whose initial goal is to create a centralized database of physically based, machine-learning, and hybrid model forecasts to enable a distributed assessment and evaluation effort. The first instance of WP-MIP focuses on global deterministic predictions using both center-specific and common initializations to facilitate sensitivity studies. Forecasts contributed by institutions across six continents will be used to develop AI-ready verification techniques that highlight the strengths and weaknesses of each class of prediction system, with the goal of establishing best-practice guidance to model developers and national weather centers. The broad engagement of the operational and forecast-evaluation communities in WP-MIP will ensure that the project results are highly relevant to the development and deployment of next-generation weather prediction systems.

arXiv:2604.16643Submitted: Jul 21, 2026v2

Machine-learning weather models now compete with established numerical weather prediction systems on some medium-range scores, but the comparison problem has not been solved. Centers initialize models differently, verify against different analyses, archive different fields, and emphasize different use cases. WP-MIP addresses that gap as infrastructure rather than as a new forecasting model: it defines a shared intercomparison project for physically based, AI-based, and hybrid prediction systems.

The paper’s central claim is pragmatic. If national centers are going to use AI weather prediction in operations, they need a common archive and verification workflow that can reveal where each model class is skillful, physically credible, and operationally useful. The first WP-MIP instance focuses on global deterministic forecasts, with later expansion paths for probabilistic and ensemble prediction.

Core Contribution

WP-MIP’s main contribution is the protocol design. The project separates forecasts into two initialization streams: an own initial conditions stream, where each center uses its operational or preferred initialization, and a same initial conditions stream, where models start from common ECMWF initial states. That split is the important design choice. OIC reflects what a forecasting center can deliver in practice; SIC removes part of the initialization confound and lets evaluators focus more directly on model behavior.

Figure 1 summarizes this structure, including the core protocol and two subprojects.

Figure 1. Project overview schematic. The core protocol (blue) is divided by initial condition (IC) specifications into the ``same initial condition'' (SIC) and ``own initial condition'' (OIC) streams, described in more detail in section~dataprotocol. The two WP-MIP subprojects (SP1 and SP2) are described in sections~dataprotocolsp1 and dataprotocolsp2, respectively.

The paper also defines WP-MIP as a distributed evaluation effort, not a single leaderboard. Forecasts are stored centrally, but verification is meant to be developed by participating groups. That matters because the open questions are broader than headline error scores: physical consistency, extremes, spatial structure, tropical phenomena, subseasonal signals, and regional forecast value all require different diagnostics.

Technical Approach

The project compares three model families: conventional physical models, AI weather prediction systems, and hybrid systems that combine data-driven components with physical-model machinery. The first phase uses global deterministic forecasts and a common 0.25° WP-MIP grid for shared diagnostics. The figures show evaluation against multiple reference analyses, with uncertainty estimated by a 1000-member bootstrap in the spectral diagnostics.

The protocol’s verification program is wider than pointwise error. The paper discusses conventional bias and RMSE for global 500 hPa temperature, spectral kinetic-energy diagnostics for 250 hPa winds at day 10, falsification tests for physically impossible states, spatial verification for feature structure and displacement errors, and event-oriented evaluation for extremes. The falsification section is especially relevant for AI systems: the authors cite negative precipitation as a clear example of an impossible atmospheric prediction, then broaden the point to conservation and process-level consistency.

Figure 7 shows the spectral amplitude-ratio diagnostic used to compare kinetic-energy structure against the analysis reference. It is a useful example of the paper’s broader argument: a model can have competitive scalar scores while still putting too much or too little variance at specific spatial scales.

Figure 7. Spectral amplitude ratio for global 250~hPa kinetic energy at day 10 (Eq.~sar) . A ratio of 1 (unity) is identified with a black dashed line. Color-coded vertical lines for hybrid-model panels (c and f) represent spectral nudging filter cutoffs.

Results and Analysis

This is not a paper that reports a finished WP-MIP ranking. Its evidence is a project archive, protocol specification, and sample diagnostics from early contributed forecasts. The 500 hPa temperature figures compare bias and RMSE across physical, AI, and hybrid models, with ECMWF’s physical forecast plotted as a reference in the AI and hybrid panels. Those plots are designed less to crown a winner than to show how OIC and SIC comparisons expose different sources of error.

The spectral figures are more diagnostic. At day 10 for global 250 hPa winds, the AI-model panels show a common loss of small-scale kinetic energy at high global wavenumber, while hybrid systems are closer to the analysis spectrum in the shown cases. Physical models vary more by center at small scales, which is expected because their numerics, diffusion, and resolution choices differ. The paper’s interpretation is careful: these diagnostics identify behavior that standard RMSE can hide, but they do not by themselves determine operational value.

The strongest part of the work is the protocol: OIC versus SIC, multi-analysis verification, common gridding, 6-hourly fields, and a planned set of phenomenon-based studies. The weaker part is timing. The archive covers a 1-year period, which the authors explicitly describe as limited for rare or extreme events. The protocol also lacks ensemble-based prediction in its first instance, even though recent AI successes suggest that future intercomparisons will need probabilistic components. For operational centers, the paper is most useful as a blueprint for how to evaluate next-generation forecast systems without reducing the question to one global score.

Evidence Box

moderate

Key Claims

  • Common archive enables fair comparison of physical, AI, and hybrid forecasts
  • Own-initial-condition and same-initial-condition streams separate operational performance from initialization sensitivity
  • AI-ready verification should test physical consistency, spatial structure, extremes, and phenomena
  • Future intercomparisons need probabilistic and ensemble components

Key Results

  • First WP-MIP instance focuses on global deterministic forecasts from institutions across six continents
  • Forecast fields are provided every 6 hours for case, feature-based, and regional studies
  • Global 500 hPa temperature bias and RMSE are evaluated against multiple reference analyses
  • Day-10 global 250 hPa wind spectra use 1000-member bootstrap confidence intervals on a 0.25° WP-MIP grid

Limitations & Caveats

  • Initial WP-MIP archive covers only a 1-year period, limiting rare-event sampling
  • First instance excludes ensemble-based probabilistic prediction
  • Preliminary diagnostics do not provide a final ranking of model classes
  • Protocol replacement may be needed as AIWP training regimes overlap the 2024 evaluation period

Related Articles

Readers are encouraged to consult the original arXiv paper for complete details. SOTA Papers does not make claims beyond what is supported by the authors' reported evidence.