ACE2-ERA5: Data-Driven Atmospheric Emulator
- The paper shows that ACE2-ERA5, a data-driven emulator, leverages a 6‑hour prediction loss to capture fast atmospheric dynamics while revealing limitations in slow, quasi‑periodic modes.
- ACE2-ERA5 is an autoregressive model based on spherical Fourier neural operators, trained on ERA5 reanalysis with boundary forcings like SST, SIC, and global CO2.
- ACE2-ERA5 benchmarks against physics‑based and hybrid models, excelling in midlatitude vertical trends but underperforming in capturing the QBO, heat extremes, and radiative responses.
Searching arXiv for ACE2-ERA5 and directly related benchmarking papers. arXiv search query: ACE2-ERA5 ACE2 NeuralGCM ERA5 benchmarking ACE2-ERA5 is a fully data-driven atmospheric emulator in the Ai2 Climate Emulator family, trained on ERA5 reanalysis to reproduce ERA5-like atmospheric evolution at 6-hour intervals. In its basic formulation, the model advances the atmospheric state according to with , where denotes prescribed boundary forcings and are network parameters learned by minimizing a 6-hour prediction loss (Baxter et al., 6 Oct 2025). Across the 2025 benchmarking literature, ACE2-ERA5 is presented less as a short-range forecast system than as a test case for whether a model optimized on fast atmospheric dynamics can also reproduce slower, physically meaningful circulation, radiative, and regional thermodynamic behaviors that matter for climate variability and out-of-distribution use (Baxter et al., 6 Oct 2025, Loon et al., 15 Feb 2025, Rucker et al., 31 Oct 2025).
1. Definition and model formulation
ACE2-ERA5 is described as an atmosphere-only, autoregressive emulator based on a spherical Fourier neural operator. It is trained on ERA5 reanalysis and then run autoregressively in a stable 37-member lagged ensemble (Baxter et al., 6 Oct 2025). In the circulation benchmark, the model’s evolution equation is written as
with training based on
The forcings include sea surface temperature, sea ice, downward solar radiation, and for ACE2-ERA5 also global-mean (Baxter et al., 6 Oct 2025).
A later regional-trend benchmark presents the same model class in closely related notation, describing ACE2 as an SFNO-based emulator trained with a global MSE loss,
again with 0 hours and forcings including SST, SIC, global mean CO1, and top-of-atmosphere shortwave flux (Rucker et al., 31 Oct 2025). This suggests that the core ACE2-ERA5 conception is stable across studies: a learned global atmosphere propagator conditioned on prescribed lower-boundary and radiative forcings.
The radiative-response study further specifies the data representation used for one pretrained ACE2-ERA5 configuration: ERA5 regridded to 1° horizontal resolution and 8 vertical levels, with training periods 1940–1995, 2011–2019, and 2021–2022, while other years are reserved for validation and testing (Loon et al., 15 Feb 2025). That paper also states that ACE2-ERA5 conserves the global mean moisture budget, but not the energy budget (Loon et al., 15 Feb 2025).
2. Training target, forcing structure, and relation to ERA5
ACE2-ERA5 is trained on ERA5 reanalysis rather than on GCM output. This distinction is central in the 2025 papers, which repeatedly define the model as learning atmospheric evolution “as represented by ERA5” (Loon et al., 15 Feb 2025). In the regional thermodynamic benchmark, ERA5 is also the observational reference: “ACE2-ERA5” denotes the ACE2 model trained on ERA5 data, while the “ERA5 trend” is the historical trend computed from the ERA5 reanalysis dataset and compared to model ensembles (Rucker et al., 31 Oct 2025).
The forcing interface is a defining part of the emulator. Across the papers, prescribed inputs include SST, sea-ice fraction or sea ice, incoming or downward solar radiation, and atmospheric or global mean CO2 (Baxter et al., 6 Oct 2025, Loon et al., 15 Feb 2025, Rucker et al., 31 Oct 2025). In this sense, ACE2-ERA5 is not a free-running coupled climate model; it is a forced atmospheric emulator whose climate behavior must be interpreted in relation to the imposed lower-boundary and radiative histories.
The choice of ERA5 as training and reference data places ACE2-ERA5 within a broader methodological landscape in which ERA5 functions as a benchmark-grade reanalysis for downstream climate and atmospheric tasks. In a wind-power data-selection study, ERA5 is described as “a reasonable reference” and is used as the ground truth-like comparator for all model evaluations (Morelli et al., 2024). In a lightning-classification study, ERA5 provides vertically resolved atmospheric inputs on model levels beyond the tropopause, forming an input layer of approximately 670 features for deep learning (Ehrensperger et al., 2022). A multi-country wind-power validation likewise treats ERA5 as the newer reanalysis benchmark relative to MERRA-2 and finds that ERA5 generally outperforms MERRA-2 in simulated wind-power quality (Gruber et al., 2020). These studies do not concern ACE2-ERA5 directly, but they clarify why an emulator trained on ERA5 can be framed as learning from a high-quality and widely used atmospheric reference rather than from an arbitrary archive.
At the same time, the radiative-response paper emphasizes a crucial limitation of this design: ERA5 does not assimilate direct observations of top-of-atmosphere radiative fluxes and has known energy-balance biases relative to observations (Loon et al., 15 Feb 2025). A plausible implication is that ACE2-ERA5 may inherit strengths in circulation realism while also inheriting deficiencies in energetics.
3. Atmospheric circulation variability benchmarks
The most direct assessment of ACE2-ERA5 as a dynamical emulator is the circulation benchmark comparing it with ERA5, AMIP physics-based models, and NeuralGCM using four diagnostics: the quasi-biennial oscillation, convectively coupled equatorial waves, extratropical eddy–mean flow interaction, and Southern Annular Mode propagation (Baxter et al., 6 Oct 2025). These diagnostics are explicitly described as “dynamical” tests probing wave–mean-flow physics rather than low-error pattern matching.
For the QBO, the paper defines the index as the monthly, latitude-weighted mean zonal wind averaged over 3 to 4 at 50 hPa. The observed value in ERA5 is 5 months (Baxter et al., 6 Oct 2025). ACE2-ERA5 reproduces a realistic amplitude range, about 39–42 m/s, closer to ERA5 than the AMIP models or NeuralGCM, but it fails to produce a regular oscillation or clear downward propagation (Baxter et al., 6 Oct 2025). The study links this deficiency to weak representation of stratosphere–troposphere coupling and gravity-wave-driven mean-flow interaction, and also notes a structural limitation: ACE2-ERA5 has very limited vertical resolution in the stratosphere, with only one vertically averaged stratospheric layer (Baxter et al., 6 Oct 2025).
For convectively coupled equatorial waves, evaluated using Wheeler–Kiladis wavenumber–frequency spectra of daily precipitation, ACE2-ERA5 captures the main spectral peaks associated with the Madden–Julian Oscillation, Kelvin waves, and equatorial Rossby waves, and improves on CESM2-WACCM in the MJO band around wavenumbers 2–5 (Baxter et al., 6 Oct 2025). It nevertheless underestimates higher-frequency Kelvin-wave power, higher-wavenumber Rossby waves, and especially the asymmetric mixed Rossby-gravity and inertio-gravity components (Baxter et al., 6 Oct 2025). The resulting profile is scale dependent: strong performance on large-scale tropical variability and weaker fidelity at the faster, smaller-scale end of the spectrum.
For extratropical eddy momentum flux co-spectra at 250 hPa, ACE2-ERA5 reproduces the broad structure of the observed eddy fluxes and correctly places them relative to the critical line 6, indicating that it captures the basic baroclinic eddy–mean flow physics (Baxter et al., 6 Oct 2025). However, it underestimates the spectral density of the eddy momentum flux compared with ERA5, while NeuralGCM performs somewhat better in amplitude (Baxter et al., 6 Oct 2025).
For the Southern Annular Mode, whose intrinsic periodicity is about 150 days in the paper’s formulation, ACE2-ERA5 does not consistently reproduce the ERA5 spectral peak across ensemble members (Baxter et al., 6 Oct 2025). It does capture the cross-EOF feedback structure reasonably well and shows some over-persistence similar to physics-based models, specifically by slightly overestimating the 7 transition and underestimating the reverse transition (Baxter et al., 6 Oct 2025).
The central conclusion of this benchmark is that ACE2-ERA5 succeeds on shorter-timescale dynamics but struggles with long-timescale, quasi-periodic phenomena (Baxter et al., 6 Oct 2025). The authors connect this directly to training on 6-hour prediction error, which primarily rewards fast dynamics and may not robustly reinforce processes whose signatures emerge over 8 months or 9 days (Baxter et al., 6 Oct 2025).
4. Regional thermodynamic trends and extremes
A second major line of evaluation concerns historical regional thermodynamic trends over 1981–2014. Here ACE2-ERA5 is benchmarked against NeuralGCM, a hybrid model with a dynamical core and learned subgrid parameterization, and a large AMIP land–atmosphere model ensemble (Rucker et al., 31 Oct 2025). The criterion for whether a model “captures” the ERA5 trend is whether the ERA5 value lies within the model’s 5–95% ensemble range (Rucker et al., 31 Oct 2025).
The benchmark computes regionally and annually averaged fields with a latitude-weighted mean and then estimates linear trends from the resulting annual time series (Rucker et al., 31 Oct 2025). It examines latitudinal mean temperature trends, vertical temperature trends, heat-extreme trends, and drying trends.
For Arctic Amplification, all models capture the broad signal, largely because SST and SIC are prescribed in the atmosphere-only setting; both AI models do at least as well as the physics-based ensemble (Rucker et al., 31 Oct 2025). For tropical upper-tropospheric warming, a known bias in physics-based models is excess warming aloft. The paper finds that both ACE2 and NeuralGCM capture the tropical upper-level warming trend better than physics-based models (Rucker et al., 31 Oct 2025). ACE2, however, does not fully capture near-surface warming over tropical ocean, which the paper relates to strong cooling in the eastern tropical Pacific and to the model’s limited number of hybrid-sigma levels below 850 hPa (Rucker et al., 31 Oct 2025).
The strongest reported success for ACE2 occurs in Northern Hemisphere extratropical vertical temperature trends. In the midlatitude vertical structure, ACE2 outperforms both NeuralGCM and the physics-based ensemble, reproducing the observed vertical temperature trend through much of the atmosphere more effectively than the alternatives (Rucker et al., 31 Oct 2025). This is presented as notable because NeuralGCM contains a dynamical core, which might have been expected to improve free-tropospheric structure (Rucker et al., 31 Oct 2025).
The benchmark’s results for heat extremes are more mixed. Heat-extreme trends are evaluated using TXx, constructed from daily maxima derived from the four daily 6-hourly samples and then aggregated to yearly maxima (Rucker et al., 31 Oct 2025). Over Western Europe, ACE2 shows similar performance to AMIP and NeuralGCM; over the Midwest US, NeuralGCM does better while ACE2 does not capture the trend satisfactorily; over the Southwest US, the AI models do not capture the observed heat-extreme trend, and NeuralGCM in particular underestimates it (Rucker et al., 31 Oct 2025). The authors interpret these failures as evidence that land representation matters and note that neither ACE2 nor NeuralGCM includes an explicit land surface model with soil moisture and related feedbacks (Rucker et al., 31 Oct 2025).
For drying trends, the paper focuses on 700-hPa specific humidity in the Southwest US and Southern South America. Physics-based AMIP models and NeuralGCM largely fail to simulate the observed drying, whereas ACE2 performs better: it gets the Southwest US drying trend closer to ERA5 and captures the drying over South America (Rucker et al., 31 Oct 2025). Even so, the paper treats drying as a difficult benchmark and again points to the absence of explicit land variables as likely important (Rucker et al., 31 Oct 2025).
5. Radiative response, Green’s functions, and causal process tests
ACE2-ERA5 has also been evaluated through Green’s function experiments designed to estimate the sensitivity of top-of-atmosphere radiative flux to localized SST perturbations (Loon et al., 15 Feb 2025). This benchmark differs from trend and circulation tests by framing evaluation as a causal physical-response problem: if a patch of SST is perturbed, does the emulator generate a plausible TOA radiative response?
The study first runs a 20-year control simulation forced with an annually repeating climatology derived from ERA5 over 1971–2020, obtaining a baseline mean global TOA radiation 0 (Loon et al., 15 Feb 2025). It then applies 109 localized SST patches, with 1 K and 2 K perturbations, each integrated for 12 years with the first 2 years discarded as spin-up (Loon et al., 15 Feb 2025). For each patch 3, the radiative anomaly is
4
The paper defines an area-normalized Green’s function for globally averaged TOA radiation 5 with respect to local SST 6 and then uses it to linearly reconstruct radiative responses to arbitrary SST anomaly patterns (Loon et al., 15 Feb 2025).
The resulting ACE2-ERA5 sensitivity map is described as qualitatively similar to the multi-model mean of five GCMs (Loon et al., 15 Feb 2025). It shows negative feedbacks in regions of deep convection, especially the tropical West Pacific, and positive feedbacks in subsidence and low-cloud regions, especially the subtropical East Pacific (Loon et al., 15 Feb 2025). The paper interprets these structures in physically familiar terms: warming in East Pacific low-cloud regions reduces low cloud cover and reflected shortwave radiation, while warming in the West Pacific warm pool can alter upper-tropospheric temperatures and stability remotely, increasing low cloud amount elsewhere and yielding negative feedbacks (Loon et al., 15 Feb 2025).
This process-level success is counterbalanced by a quantitative failure. When the ACE2-ERA5 Green’s function is convolved with historical SST anomalies, the reconstructed TOA radiative response does not show the expected declining trend and fails to become sufficiently more negative over time (Loon et al., 15 Feb 2025). Variability is also too small compared with observed TOA imbalance variability, and even when driven by historical ERA5 forcing, ACE2-ERA5 does not reproduce yearly TOA radiation correctly, including during the training period (Loon et al., 15 Feb 2025).
The paper offers several explanations: ERA5 has energy-budget biases and does not assimilate direct TOA radiative flux observations; ACE2-ERA5 conserves moisture but not energy; and the SST patch experiments are out-of-distribution because the model was trained on historical trajectories rather than artificial isolated patch perturbations (Loon et al., 15 Feb 2025). The broader inference is explicit: qualitative process realism does not guarantee quantitatively correct forced-response prediction.
6. Interpretation, limitations, and significance for out-of-distribution applications
Across the 2025 evaluations, ACE2-ERA5 is characterized by a mixed but coherent profile. It can reproduce substantial atmospheric dynamics learned from reanalysis alone, including the spectra of large-scale tropical waves, the broad geometry of extratropical eddy–jet interactions, Arctic warming, tropical upper-tropospheric warming, and much of the midlatitude vertical temperature-trend structure (Baxter et al., 6 Oct 2025, Rucker et al., 31 Oct 2025). It also generates a physically sensible SST-to-TOA-radiation sensitivity pattern in Green’s function experiments (Loon et al., 15 Feb 2025).
Its principal weaknesses are equally consistent across papers. ACE2-ERA5 struggles with slow quasi-periodic modes such as the QBO’s 7-month cycle and the SAM’s 8-day propagation (Baxter et al., 6 Oct 2025). It does not capture regional heat-extreme trends over the US Southwest and does not uniformly represent lower-tropospheric tropical ocean trends (Rucker et al., 31 Oct 2025). In radiative-response tests, it likely underestimates the response to historical warming and fails to reproduce yearly TOA radiation correctly (Loon et al., 15 Feb 2025).
Several limitations recur. One is optimization mismatch: training on 6-hour prediction error strongly rewards fast dynamics, so slow oscillatory mechanisms may contribute too weakly to the loss to be learned robustly (Baxter et al., 6 Oct 2025). A second is structural representation, especially the very limited stratospheric vertical resolution noted in the QBO analysis and the sparse low-level vertical structure implicated in tropical ocean trend errors (Baxter et al., 6 Oct 2025, Rucker et al., 31 Oct 2025). A third is missing physical closure: the radiative-response study explicitly identifies lack of energy conservation as a likely contributor to deficient TOA-radiation behavior (Loon et al., 15 Feb 2025). A fourth is land-process omission: failures in heat extremes are interpreted in terms of missing explicit land-surface and soil-moisture feedbacks (Rucker et al., 31 Oct 2025).
These limitations matter most for out-of-distribution use. The circulation paper states that the four diagnostics are an initial credibility test for AI climate emulators because climate response depends on whether the model captures underlying circulation dynamics rather than merely historical climatology (Baxter et al., 6 Oct 2025). The radiative-response study reaches a parallel conclusion through Green’s functions: the emulator can encode physically meaningful causal relationships yet still fail on out-of-distribution predictive tasks (Loon et al., 15 Feb 2025). Taken together, these results position ACE2-ERA5 as a system with substantial learned dynamical structure but scale-dependent physical fidelity.
7. Position within AI climate modeling
ACE2-ERA5 is repeatedly compared with two alternatives: physics-based AMIP atmosphere-land models and NeuralGCM, a hybrid model combining a dynamical core with learned parameterization (Baxter et al., 6 Oct 2025, Rucker et al., 31 Oct 2025). This comparison clarifies its methodological role. Relative to AMIP, ACE2-ERA5 is often comparable to or better than the physics-based ensemble on several circulation and thermodynamic diagnostics, particularly tropical wave spectra, aspects of extratropical eddy–mean flow interaction, tropical upper-tropospheric warming, and extratropical vertical temperature trends (Baxter et al., 6 Oct 2025, Rucker et al., 31 Oct 2025). Relative to NeuralGCM, the pattern is mixed: NeuralGCM performs somewhat better for extratropical eddy momentum-flux amplitude and captures some heat-extreme benchmarks better, while ACE2 performs better in the QBO amplitude range, in parts of the drying benchmark, and especially in midlatitude vertical temperature trends (Baxter et al., 6 Oct 2025, Rucker et al., 31 Oct 2025).
The broader significance assigned to ACE2-ERA5 in these papers is therefore not that it is uniformly superior, nor that it is unphysical. The consistent conclusion is instead that a fully data-driven emulator trained on ERA5 can learn many dynamical and thermodynamic structures that are relevant to climate, while still missing slow modes, land-coupled extremes, and quantitative forced radiative responses that are crucial for robust climate projection (Baxter et al., 6 Oct 2025, Loon et al., 15 Feb 2025, Rucker et al., 31 Oct 2025). This suggests an emerging evaluation principle for AI climate emulators: short-horizon skill and realistic climatological patterning are insufficient on their own; benchmarking must also test wave–mean-flow dynamics, multi-decadal regional trends, radiative causal structure, and other metrics tied to physical credibility under changed forcings.