---
title: 'ACE2-ERA5: Data-Driven Atmospheric Emulator'
url: https://www.emergentmind.com/topics/ace2-era5
type: topic
---

# ACE2-ERA5: Data-Driven Atmospheric Emulator

Searching arXiv for ACE2-ERA5 and directly related benchmarking papers.
arXiv search query: ACE2-ERA5 ACE2 NeuralGCM ERA5 benchmarking
ACE2-ERA5 is a fully data-driven atmospheric emulator in the Ai2 Climate Emulator family, trained on ERA5 reanalysis to reproduce ERA5-like atmospheric evolution at 6-hour intervals. In its basic formulation, the model advances the atmospheric state \(x(t)\) according to \(x(t+\Delta t)=N(x(t),b(t),\theta)\) with \(\Delta t=6\ \text{hours}\), where \(b(t)\) denotes prescribed boundary forcings and \(\theta\) are network parameters learned by minimizing a 6-hour prediction loss [2510.04466]. Across the 2025 benchmarking literature, ACE2-ERA5 is presented less as a short-range forecast system than as a test case for whether a model optimized on fast atmospheric dynamics can also reproduce slower, physically meaningful circulation, radiative, and regional thermodynamic behaviors that matter for climate variability and out-of-distribution use [2510.04466], [2502.10893], [2511.00274].

## 1. Definition and model formulation

ACE2-ERA5 is described as an atmosphere-only, autoregressive emulator based on a spherical Fourier neural operator. It is trained on ERA5 reanalysis and then run autoregressively in a stable 37-member lagged ensemble [2510.04466]. In the circulation benchmark, the model’s evolution equation is written as
\[
x(t+\Delta t)=N(x(t),b(t),\theta), \qquad \Delta t=6\ \text{hours},
\]
with training based on
\[
\mathcal{L}= \| x(t+\Delta t)-N(x(t),b(t),\theta)\|_2 .
\]
The forcings \(b(t)\) include sea surface temperature, sea ice, downward solar radiation, and for ACE2-ERA5 also global-mean \(\mathrm{CO}_2\) [2510.04466].

A later regional-trend benchmark presents the same model class in closely related notation, describing ACE2 as an SFNO-based emulator trained with a global MSE loss,
\[
\mathcal{L}= \frac{1}{N}\sum^N_{i=1} \left\|x_{i}(t+\Delta t) - \mathcal{N}(x_{i}(t),b_i(t),\theta)\right\|_2^2,
\]
again with \(\Delta t = 6\) hours and forcings including SST, SIC, global mean CO\(_2\), and top-of-atmosphere shortwave flux [2511.00274]. This suggests that the core ACE2-ERA5 conception is stable across studies: a learned global atmosphere propagator conditioned on prescribed lower-boundary and radiative forcings.

The radiative-response study further specifies the data representation used for one pretrained ACE2-ERA5 configuration: ERA5 regridded to 1° horizontal resolution and 8 vertical levels, with training periods 1940–1995, 2011–2019, and 2021–2022, while other years are reserved for validation and testing [2502.10893]. That paper also states that ACE2-ERA5 conserves the global mean moisture budget, but not the energy budget [2502.10893].

## 2. Training target, forcing structure, and relation to ERA5

ACE2-ERA5 is trained on ERA5 reanalysis rather than on GCM output. This distinction is central in the 2025 papers, which repeatedly define the model as learning atmospheric evolution “as represented by ERA5” [2502.10893]. In the regional thermodynamic benchmark, ERA5 is also the observational reference: “ACE2-ERA5” denotes the ACE2 model trained on ERA5 data, while the “ERA5 trend” is the historical trend computed from the ERA5 reanalysis dataset and compared to model ensembles [2511.00274].

The forcing interface is a defining part of the emulator. Across the papers, prescribed inputs include SST, sea-ice fraction or sea ice, incoming or downward solar radiation, and atmospheric or global mean CO\(_2\) [2510.04466], [2502.10893], [2511.00274]. In this sense, ACE2-ERA5 is not a free-running coupled climate model; it is a forced atmospheric emulator whose climate behavior must be interpreted in relation to the imposed lower-boundary and radiative histories.

The choice of ERA5 as training and reference data places ACE2-ERA5 within a broader methodological landscape in which ERA5 functions as a benchmark-grade reanalysis for downstream climate and atmospheric tasks. In a wind-power data-selection study, ERA5 is described as “a reasonable reference” and is used as the ground truth-like comparator for all model evaluations [2411.11630]. In a lightning-classification study, ERA5 provides vertically resolved atmospheric inputs on model levels beyond the tropopause, forming an input layer of approximately 670 features for deep learning [2210.11529]. A multi-country wind-power validation likewise treats ERA5 as the newer reanalysis benchmark relative to MERRA-2 and finds that ERA5 generally outperforms MERRA-2 in simulated wind-power quality [2012.05648]. These studies do not concern ACE2-ERA5 directly, but they clarify why an emulator trained on ERA5 can be framed as learning from a high-quality and widely used atmospheric reference rather than from an arbitrary archive.

At the same time, the radiative-response paper emphasizes a crucial limitation of this design: ERA5 does not assimilate direct observations of top-of-atmosphere radiative fluxes and has known energy-balance biases relative to observations [2502.10893]. A plausible implication is that ACE2-ERA5 may inherit strengths in circulation realism while also inheriting deficiencies in energetics.

## 3. Atmospheric circulation variability benchmarks

The most direct assessment of ACE2-ERA5 as a dynamical emulator is the circulation benchmark comparing it with ERA5, AMIP physics-based models, and NeuralGCM using four diagnostics: the quasi-biennial oscillation, convectively coupled equatorial waves, extratropical eddy–mean flow interaction, and Southern Annular Mode propagation [2510.04466]. These diagnostics are explicitly described as “dynamical” tests probing wave–mean-flow physics rather than low-error pattern matching.

For the QBO, the paper defines the index as the monthly, latitude-weighted mean zonal wind averaged over \(10^\circ\mathrm{S}\) to \(10^\circ\mathrm{N}\) at 50 hPa. The observed value in ERA5 is \(\sim 28\) months [2510.04466]. ACE2-ERA5 reproduces a realistic amplitude range, about 39–42 m/s, closer to ERA5 than the AMIP models or NeuralGCM, but it fails to produce a regular oscillation or clear downward propagation [2510.04466]. The study links this deficiency to weak representation of stratosphere–troposphere coupling and gravity-wave-driven mean-flow interaction, and also notes a structural limitation: ACE2-ERA5 has very limited vertical resolution in the stratosphere, with only one vertically averaged stratospheric layer [2510.04466].

For convectively coupled equatorial waves, evaluated using Wheeler–Kiladis wavenumber–frequency spectra of daily precipitation, ACE2-ERA5 captures the main spectral peaks associated with the Madden–Julian Oscillation, Kelvin waves, and equatorial Rossby waves, and improves on CESM2-WACCM in the MJO band around wavenumbers 2–5 [2510.04466]. It nevertheless underestimates higher-frequency Kelvin-wave power, higher-wavenumber Rossby waves, and especially the asymmetric mixed Rossby-gravity and inertio-gravity components [2510.04466]. The resulting profile is scale dependent: strong performance on large-scale tropical variability and weaker fidelity at the faster, smaller-scale end of the spectrum.

For extratropical eddy momentum flux co-spectra at 250 hPa, ACE2-ERA5 reproduces the broad structure of the observed eddy fluxes and correctly places them relative to the critical line \(\overline{u}=c\), indicating that it captures the basic baroclinic eddy–mean flow physics [2510.04466]. However, it underestimates the spectral density of the eddy momentum flux compared with ERA5, while NeuralGCM performs somewhat better in amplitude [2510.04466].

For the Southern Annular Mode, whose intrinsic periodicity is about 150 days in the paper’s formulation, ACE2-ERA5 does not consistently reproduce the ERA5 spectral peak across ensemble members [2510.04466]. It does capture the cross-EOF feedback structure reasonably well and shows some over-persistence similar to physics-based models, specifically by slightly overestimating the \(z_2\to z_1\) transition and underestimating the reverse transition [2510.04466].

The central conclusion of this benchmark is that ACE2-ERA5 succeeds on shorter-timescale dynamics but struggles with long-timescale, quasi-periodic phenomena [2510.04466]. The authors connect this directly to training on 6-hour prediction error, which primarily rewards fast dynamics and may not robustly reinforce processes whose signatures emerge over \(\sim 28\) months or \(\sim 150\) days [2510.04466].

## 4. Regional thermodynamic trends and extremes

A second major line of evaluation concerns historical regional thermodynamic trends over 1981–2014. Here ACE2-ERA5 is benchmarked against NeuralGCM, a hybrid model with a dynamical core and learned subgrid parameterization, and a large AMIP land–atmosphere model ensemble [2511.00274]. The criterion for whether a model “captures” the ERA5 trend is whether the ERA5 value lies within the model’s 5–95% ensemble range [2511.00274].

The benchmark computes regionally and annually averaged fields with a latitude-weighted mean and then estimates linear trends from the resulting annual time series [2511.00274]. It examines latitudinal mean temperature trends, vertical temperature trends, heat-extreme trends, and drying trends.

For Arctic Amplification, all models capture the broad signal, largely because SST and SIC are prescribed in the atmosphere-only setting; both AI models do at least as well as the physics-based ensemble [2511.00274]. For tropical upper-tropospheric warming, a known bias in physics-based models is excess warming aloft. The paper finds that both ACE2 and NeuralGCM capture the tropical upper-level warming trend better than physics-based models [2511.00274]. ACE2, however, does not fully capture near-surface warming over tropical ocean, which the paper relates to strong cooling in the eastern tropical Pacific and to the model’s limited number of hybrid-sigma levels below 850 hPa [2511.00274].

The strongest reported success for ACE2 occurs in Northern Hemisphere extratropical vertical temperature trends. In the midlatitude vertical structure, ACE2 outperforms both NeuralGCM and the physics-based ensemble, reproducing the observed vertical temperature trend through much of the atmosphere more effectively than the alternatives [2511.00274]. This is presented as notable because NeuralGCM contains a dynamical core, which might have been expected to improve free-tropospheric structure [2511.00274].

The benchmark’s results for heat extremes are more mixed. Heat-extreme trends are evaluated using **TXx**, constructed from daily maxima derived from the four daily 6-hourly samples and then aggregated to yearly maxima [2511.00274]. Over Western Europe, ACE2 shows similar performance to AMIP and NeuralGCM; over the Midwest US, NeuralGCM does better while ACE2 does not capture the trend satisfactorily; over the Southwest US, the AI models do not capture the observed heat-extreme trend, and NeuralGCM in particular underestimates it [2511.00274]. The authors interpret these failures as evidence that land representation matters and note that neither ACE2 nor NeuralGCM includes an explicit land surface model with soil moisture and related feedbacks [2511.00274].

For drying trends, the paper focuses on 700-hPa specific humidity in the Southwest US and Southern South America. Physics-based AMIP models and NeuralGCM largely fail to simulate the observed drying, whereas ACE2 performs better: it gets the Southwest US drying trend closer to ERA5 and captures the drying over South America [2511.00274]. Even so, the paper treats drying as a difficult benchmark and again points to the absence of explicit land variables as likely important [2511.00274].

## 5. Radiative response, Green’s functions, and causal process tests

ACE2-ERA5 has also been evaluated through Green’s function experiments designed to estimate the sensitivity of top-of-atmosphere radiative flux to localized SST perturbations [2502.10893]. This benchmark differs from trend and circulation tests by framing evaluation as a causal physical-response problem: if a patch of SST is perturbed, does the emulator generate a plausible TOA radiative response?

The study first runs a 20-year control simulation forced with an annually repeating climatology derived from ERA5 over 1971–2020, obtaining a baseline mean global TOA radiation \(R_c\) [2502.10893]. It then applies 109 localized SST patches, with \(+2\) K and \(-2\) K perturbations, each integrated for 12 years with the first 2 years discarded as spin-up [2502.10893]. For each patch \(p\), the radiative anomaly is
\[
\Delta_p R = R_p - R_c.
\]
The paper defines an area-normalized Green’s function for globally averaged TOA radiation \(R\) with respect to local SST \(T_j\) and then uses it to linearly reconstruct radiative responses to arbitrary SST anomaly patterns [2502.10893].

The resulting ACE2-ERA5 sensitivity map is described as qualitatively similar to the multi-model mean of five GCMs [2502.10893]. It shows negative feedbacks in regions of deep convection, especially the tropical West Pacific, and positive feedbacks in subsidence and low-cloud regions, especially the subtropical East Pacific [2502.10893]. The paper interprets these structures in physically familiar terms: warming in East Pacific low-cloud regions reduces low cloud cover and reflected shortwave radiation, while warming in the West Pacific warm pool can alter upper-tropospheric temperatures and stability remotely, increasing low cloud amount elsewhere and yielding negative feedbacks [2502.10893].

This process-level success is counterbalanced by a quantitative failure. When the ACE2-ERA5 Green’s function is convolved with historical SST anomalies, the reconstructed TOA radiative response does not show the expected declining trend and fails to become sufficiently more negative over time [2502.10893]. Variability is also too small compared with observed TOA imbalance variability, and even when driven by historical ERA5 forcing, ACE2-ERA5 does not reproduce yearly TOA radiation correctly, including during the training period [2502.10893].

The paper offers several explanations: ERA5 has energy-budget biases and does not assimilate direct TOA radiative flux observations; ACE2-ERA5 conserves moisture but not energy; and the SST patch experiments are out-of-distribution because the model was trained on historical trajectories rather than artificial isolated patch perturbations [2502.10893]. The broader inference is explicit: qualitative process realism does not guarantee quantitatively correct forced-response prediction.

## 6. Interpretation, limitations, and significance for out-of-distribution applications

Across the 2025 evaluations, ACE2-ERA5 is characterized by a mixed but coherent profile. It can reproduce substantial atmospheric dynamics learned from reanalysis alone, including the spectra of large-scale tropical waves, the broad geometry of extratropical eddy–jet interactions, Arctic warming, tropical upper-tropospheric warming, and much of the midlatitude vertical temperature-trend structure [2510.04466], [2511.00274]. It also generates a physically sensible SST-to-TOA-radiation sensitivity pattern in Green’s function experiments [2502.10893].

Its principal weaknesses are equally consistent across papers. ACE2-ERA5 struggles with slow quasi-periodic modes such as the QBO’s \(\sim 28\)-month cycle and the SAM’s \(\sim 150\)-day propagation [2510.04466]. It does not capture regional heat-extreme trends over the US Southwest and does not uniformly represent lower-tropospheric tropical ocean trends [2511.00274]. In radiative-response tests, it likely underestimates the response to historical warming and fails to reproduce yearly TOA radiation correctly [2502.10893].

Several limitations recur. One is optimization mismatch: training on 6-hour prediction error strongly rewards fast dynamics, so slow oscillatory mechanisms may contribute too weakly to the loss to be learned robustly [2510.04466]. A second is structural representation, especially the very limited stratospheric vertical resolution noted in the QBO analysis and the sparse low-level vertical structure implicated in tropical ocean trend errors [2510.04466], [2511.00274]. A third is missing physical closure: the radiative-response study explicitly identifies lack of energy conservation as a likely contributor to deficient TOA-radiation behavior [2502.10893]. A fourth is land-process omission: failures in heat extremes are interpreted in terms of missing explicit land-surface and soil-moisture feedbacks [2511.00274].

These limitations matter most for out-of-distribution use. The circulation paper states that the four diagnostics are an initial credibility test for AI climate emulators because climate response depends on whether the model captures underlying circulation dynamics rather than merely historical climatology [2510.04466]. The radiative-response study reaches a parallel conclusion through Green’s functions: the emulator can encode physically meaningful causal relationships yet still fail on out-of-distribution predictive tasks [2502.10893]. Taken together, these results position ACE2-ERA5 as a system with substantial learned dynamical structure but scale-dependent physical fidelity.

## 7. Position within AI climate modeling

ACE2-ERA5 is repeatedly compared with two alternatives: physics-based AMIP atmosphere-land models and NeuralGCM, a hybrid model combining a dynamical core with learned parameterization [2510.04466], [2511.00274]. This comparison clarifies its methodological role. Relative to AMIP, ACE2-ERA5 is often comparable to or better than the physics-based ensemble on several circulation and thermodynamic diagnostics, particularly tropical wave spectra, aspects of extratropical eddy–mean flow interaction, tropical upper-tropospheric warming, and extratropical vertical temperature trends [2510.04466], [2511.00274]. Relative to NeuralGCM, the pattern is mixed: NeuralGCM performs somewhat better for extratropical eddy momentum-flux amplitude and captures some heat-extreme benchmarks better, while ACE2 performs better in the QBO amplitude range, in parts of the drying benchmark, and especially in midlatitude vertical temperature trends [2510.04466], [2511.00274].

The broader significance assigned to ACE2-ERA5 in these papers is therefore not that it is uniformly superior, nor that it is unphysical. The consistent conclusion is instead that a fully data-driven emulator trained on ERA5 can learn many dynamical and thermodynamic structures that are relevant to climate, while still missing slow modes, land-coupled extremes, and quantitative forced radiative responses that are crucial for robust climate projection [2510.04466], [2502.10893], [2511.00274]. This suggests an emerging evaluation principle for AI climate emulators: short-horizon skill and realistic climatological patterning are insufficient on their own; benchmarking must also test wave–mean-flow dynamics, multi-decadal regional trends, radiative causal structure, and other metrics tied to physical credibility under changed forcings.

Source: https://www.emergentmind.com/topics/ace2-era5