ACE2: AI Climate Emulator v2
- ACE2 is a 450M-parameter autoregressive atmospheric emulator featuring an SFNO backbone that simulates weather-to-decadal climate variability under prescribed forcings.
- It enforces global dry air mass and moisture conservation to ensure long-term stability while reproducing phenomena such as tropical cyclones, ENSO responses, and sudden stratospheric warmings.
- The model bridges weather-scale predictions with climate-scale forced responses, serving as a versatile platform for coupled ocean simulations and stochastic downscaling applications.
Ai2 Climate Emulator version 2 (ACE2) is a 450M-parameter autoregressive machine-learning atmospheric emulator developed by the Allen Institute for Artificial Intelligence for climate-scale simulation under prescribed external boundary conditions, notably sea surface temperature (SST) and global-mean CO (Watt-Meyer et al., 2024). It operates at 6-hour temporal resolution, 1 horizontal resolution, and eight vertical layers, is formulated to exactly conserve global dry air mass and moisture, and can be stepped forward stably for arbitrarily many steps; reported throughput is about 1500 simulated years per wall clock day on 1 H100 GPU (Watt-Meyer et al., 2024). ACE2 is intended to bridge weather-scale autoregressive prediction and climate-scale forced response, reproducing variability from days to decades while exposing strengths and limitations of data-driven atmospheric emulation under changing climate conditions (Watt-Meyer et al., 2024).
1. Developmental context and model lineage
ACE2 emerged from the earlier ACE system, a 200M-parameter autoregressive emulator of an existing comprehensive 100-km-resolution global atmospheric model that used a Spherical Fourier Neural Operator (SFNO), ran stably for 100 years, nearly conserved column moisture without explicit constraints, and reproduced the reference model’s climate while requiring nearly 100x less wall clock time and being 100x more energy efficient than the reference model (Watt-Meyer et al., 2023). ACE v1 demonstrated that ML-based atmospheric emulation could be long-term stable and physically coherent, but its training and forcing formulation were not designed to assess sensitivity to varying external boundary conditions in the way climate-change applications require (Watt-Meyer et al., 2023).
ACE2 was introduced specifically to address this gap. In contrast to weather-oriented ML systems that emphasize short-horizon predictive skill, ACE2 was formulated to enable assessment of atmospheric response to varying prescribed SST and greenhouse-gas forcing over the past 80 years, while retaining subseasonal-to-decadal variability and long-run stability (Watt-Meyer et al., 2024). This developmental shift is central to its scientific identity: ACE2 is not merely a faster forecaster, but an atmospheric emulator intended to reproduce both internal variability and forced response under AMIP-like boundary forcing (Watt-Meyer et al., 2024).
Subsequent work treated ACE2 not as a single frozen model but as a platform. Variants include ACE2-SOM, which couples ACE2 to a slab ocean and trains on equilibrium climates with varying CO (Clark et al., 2024); ACE2.1-ERA5, the AIMIP Phase 1 submission adapted to a common intercomparison protocol without CO as an input (Henn et al., 7 May 2026); ACE2S, a stochastic extension used within HiRO-ACE (Perkins et al., 20 Dec 2025); and later ACE-family systems trained with “random-CO” reference simulations and an energy-conservation constraint to disentangle SST and CO effects more effectively (Clark et al., 6 Jun 2026). Taken together, these variants define a research program centered on atmospheric emulation, forcing sensitivity, and physical consistency.
2. Numerical formulation, forcings, and conservation structure
ACE2 is a spherical neural-operator model with an SFNO backbone, embedding dimension 384, eight spectral layers, and a deep nonlocal operator stack (Watt-Meyer et al., 2024). It is trained autoregressively to predict two 6-hour steps ahead, but is evaluated over rollouts ranging from days to millennia (Watt-Meyer et al., 2024). Its forcing inputs include spatially varying SST, global-mean CO, surface type fractions, incoming solar radiation, and terrain height; compared with earlier ACE configurations, explicit CO input is a defining addition (Watt-Meyer et al., 2024). A 4 version was also trained and evaluated for comparison, but the principal ACE2 configuration is 1 with eight terrain-following hybrid sigma-pressure layers (Watt-Meyer et al., 2024).
A distinctive feature of ACE2 is that conservation of global dry air mass and total atmospheric moisture is hard-wired rather than merely regularized in the loss. The dry-air constraint is expressed as
0
with
1
where 2 is surface pressure, 3 gravity, and 4 total water path (Watt-Meyer et al., 2024). Global column moisture is constrained through
5
with 6 evaporation, 7 precipitation, and an advective tendency term (Watt-Meyer et al., 2024). ACE2 applies conservation corrections as part of a corrector module after the SFNO prediction and before evaluating loss, and also imposes positivity constraints by setting negative moisture, precipitation, and radiative fluxes to zero (Watt-Meyer et al., 2024).
These design choices matter because they directly target long-horizon stability. The model is reported to remain stable for more than 1000 simulated years without drift under repeating climatological boundary conditions (Watt-Meyer et al., 2024). At the same time, ACE2 does not exactly conserve atmospheric energy, and this omission later became central in diagnosing its out-of-distribution failures under abrupt or disentangled forcing changes (Watt-Meyer et al., 2024).
3. Simulated variability, emergent phenomena, and benchmarked skill
ACE2 was presented as a model that learns atmospheric variability across a broad spectral range, from weather to decadal climate response. In the primary ACE2 study, it generated emergent phenomena including tropical cyclones, the Madden–Julian Oscillation, and sudden stratospheric warmings, and it reproduced atmospheric responses to El Niño variability as well as global temperature trends over the past 80 years (Watt-Meyer et al., 2024). Reported evaluation metrics include an 8 of 0.93 for global, annually averaged 2-m temperature in ACE2-ERA5 relative to the reference, compared with 0.97 for SHiELD versus SHiELD, and 10-year climate-mean errors that were 1.1–1.5 times reference-model ensemble variability (Watt-Meyer et al., 2024).
More targeted benchmarking refined this picture. In an assessment of regional thermodynamic trends over 1981–2014, ACE2 matched ERA5 in Arctic amplification, outperformed other models in capturing vertical temperature trends in the midlatitudes, and generally performed better than physics-based land-atmosphere models on several regional thermodynamic diagnostics (Rucker et al., 31 Oct 2025). However, it did not capture regional trends in heat extremes over the US Southwest, and it did not capture drying trends in arid regions consistently, although it captured drying in South America and came closer than the comparison models in the US Southwest (Rucker et al., 31 Oct 2025). The same study reported that ACE2 underestimates interannual TXx variability across all evaluated regions (Rucker et al., 31 Oct 2025).
A complementary benchmark of atmospheric circulation variability focused on four dynamical diagnostics: the quasi-biennial oscillation (QBO), equatorial wavenumber-frequency spectra, extratropical eddy momentum-flux co-spectra, and Southern annular mode (SAM) propagation (Baxter et al., 6 Oct 2025). ACE2 captured the spectra of large-scale tropical waves and extratropical eddy–mean flow interactions, including critical levels, and closely reproduced large-scale tropical convectively coupled wave structure (Baxter et al., 6 Oct 2025). Yet it failed to reproduce a realistic 9-month QBO and did not recover the observed 0-day SAM spectral peak, even though its EOF-based correlation structure for SAM compared favorably with physics-based models (Baxter et al., 6 Oct 2025). The documented interpretation was that fast-timescale 6-hourly training and coarse stratospheric representation favor weather-scale variability while underspecifying slow, oscillatory modes (Baxter et al., 6 Oct 2025).
The resulting picture is asymmetrical rather than uniformly positive or negative. ACE2 is demonstrably strong for many day-to-seasonal and interannual dynamical structures, including ENSO-related regression patterns, tropical wave spectra, and extratropical eddy critical-layer behavior, but it is weaker for slow stratospheric oscillations, some land-controlled thermodynamic trends, and certain extreme-event statistics (Watt-Meyer et al., 2024, Baxter et al., 6 Oct 2025, Rucker et al., 31 Oct 2025).
4. Bias structure, out-of-sample behavior, and stress tests
A major line of research on ACE2 concerns out-of-sample generalization under warming climates. An analysis of boreal winter land temperatures for 1996–2010 found that ACE2, when simulating this out-of-sample period, exhibits a cold bias relative to ERA5 despite inclusion of explicit CO1 forcing, with a global mean temperature bias of 2 K (Landsberg et al., 26 Sep 2025). The mean climate predicted by ACE2 for 1996–2010 resembles the observed climate of 15–20 years earlier, with lags up to 3 years in some regions such as the Eastern U.S. (Landsberg et al., 26 Sep 2025). The cold bias is especially large over North America, Europe, and Russia; it is strongest in the coldest 10% of predicted temperatures, particularly winter cold extremes, and is smaller in boreal summer (Landsberg et al., 26 Sep 2025). The study further noted that ACE2 bias patterns align with regions, seasons, and parts of the temperature distribution that have experienced the most intense historical warming, and argued that simply including forcings such as CO4 does not completely resolve training-set anchoring (Landsberg et al., 26 Sep 2025).
The AIMIP Phase 1 intercomparison exposed a related sensitivity to protocol. Under AIMIP constraints, ACE2.1-ERA5 used only ERA5 data from 1979–2014 and was not allowed to use CO5 as an input forcing (Henn et al., 7 May 2026). In that setting, ACE2.1-ERA5 was among the better AI models for mean biases and ENSO response relative to ERA5, but it underpredicted global warming in both train and test periods and showed severe limitations in out-of-sample 6 K and 7 K SST experiments, including implausible cooling over land (Henn et al., 7 May 2026). This protocol dependence is important because it demonstrates that mean-state fidelity and ENSO skill do not guarantee robust forced-response extrapolation.
Process-based Green’s-function tests produced a similarly mixed assessment. In reanalysis-based experiments, ACE2-ERA5 generated a sensitivity map of top-of-atmosphere radiative response to SST perturbations that was qualitatively consistent with physical expectations and cloud-feedback theory, but likely underestimated the radiative response to historical warming and failed to capture the expected increasing negative radiative response with warming in linear historical reconstructions (Loon et al., 15 Feb 2025). By contrast, when ACE2 was trained to emulate the E3SMv3 atmospheric model and evaluated using the Green’s Function Model Intercomparison Project protocol, the spatial patterns of top-of-atmosphere radiative response were qualitatively similar to the reference model, the area-weighted spatial pattern correlation was 0.53, and the full GFMIP suite could be completed in 2.3 wall-clock days on a single NVIDIA A100 GPU versus 331 days for EAMv3 on eight HPC nodes (Wu et al., 13 May 2025). The same study reported statistically significant discrepancies for some SST patches, especially over the subtropical northeast Pacific, and attributed these primarily to insufficient diversity in SST patterns sampled during training (Wu et al., 13 May 2025).
A common misconception is that adding a greenhouse-gas input or achieving strong in-sample climate skill is sufficient for climate-change robustness. The ACE2 literature rejects that simplification: explicit CO8 forcing, exact moisture and dry-air conservation, and strong historical-skill metrics each help, but none by themselves eliminate out-of-sample bias or disentangle causal responses to SST and CO9 (Watt-Meyer et al., 2024, Landsberg et al., 26 Sep 2025, Loon et al., 15 Feb 2025).
5. Coupled systems and derived ACE2 variants
ACE2 became the basis for several distinct coupled and stochastic systems, each probing a different aspect of climate emulation.
| System | Configuration | Principal finding |
|---|---|---|
| ACE2-SOM | ACE2 coupled to a slab ocean | Strong equilibrium skill; non-equilibrium jumps and energy non-conservation |
| ACE2-NEMO | ACE2 coupled to full-depth NEMO ocean | Stable 70-year runs; weak ENSO amplitude and biased forced response |
| ACE2S / HiRO-ACE | Stochastic ACE2 with diffusion downscaling | Restores coarse-grid variability and supports 3 km precipitation downscaling |
| Random-CO0 ACE2 | ACE-family model trained with independent SST and CO1 variation plus energy constraint | Improves disentangling of SST and CO2 effects and difficult forcing generalization |
ACE2-SOM coupled ACE2 to a differentiable slab ocean model and trained on equilibrium SHiELD-SOM climates with 1x, 2x, and 4x CO3, holding out 3x CO4 and ramp scenarios for out-of-sample tests (Clark et al., 2024). It reduced global RMS error in time-mean surface temperature by 74% and precipitation by 70% versus a coarser physics-based baseline, and reduced RMSE of the time and ensemble mean of all variables by 54–96% across equilibrium climates (Clark et al., 2024). It also reproduced vertical warming structure and changes in extreme precipitation up to the 99.9999th percentile (Clark et al., 2024). Yet under non-equilibrium forcing it showed unphysical regime shifts in stratospheric fields during ramped CO5 increase and jumped unrealistically quickly to the 4xCO6 state after abrupt quadrupling, violating global energy conservation and learning unphysical flux sensitivities (Clark et al., 2024).
A later ACE-family advance directly targeted this failure mode by introducing “random-CO7” reference simulations in which SST and CO8 vary independently, combined with a total-energy conservation constraint (Clark et al., 6 Jun 2026). Trained on a balance of AMIP, equilibrium-climate, and random-CO9 data, this model accurately emulated scenarios in which earlier ACE variants had failed, including AMIP 0 K and slab-ocean-coupled abrupt 4xCO1, while retaining skill in scenarios where previous models already performed well (Clark et al., 6 Jun 2026). The paper’s stated interpretation was that prior ACE models had been limited by correlated SST–CO2 training data that prevented learning their separate effects (Clark et al., 6 Jun 2026).
ACE2-NEMO extended the platform into fully coupled ocean modeling by interactively coupling ACE2 to the full-depth NEMO ocean model in 70-year historical and control simulations (Antonio et al., 30 Mar 2026). These experiments were described as the first multi-decadal integrations of a machine-learned atmosphere interacting with a full-depth dynamical ocean (Antonio et al., 30 Mar 2026). The coupled system produced realistic fast-timescale air–sea coupling in the tropical Pacific and remained stable without major drifts, but El Niño-like variability had very low amplitude and appeared close to red noise because atmospheric feedback in the tropical Pacific was weak (Antonio et al., 30 Mar 2026). Its historical CO3-forced response initially agreed with EC-Earth3P but later deviated because ACE2 produced reduced downward short-wave radiation (Antonio et al., 30 Mar 2026).
A separate branch pursued stochasticity and high-resolution downscaling. In HiRO-ACE, ACE2S is a stochastic extension of deterministic ACE2 that replaces deterministic instance normalization with conditional layer normalization and injects 64 isotropic Gaussian white-noise channels, enabling autoregressive sampling from the conditional distribution of atmospheric states (Perkins et al., 20 Dec 2025). After pretraining on 44 years of ERA5 and fine-tuning on coarsened output from the 3 km global storm-resolving model X-SHiELD, ACE2S achieved a time-mean RMS precipitation bias of 0.23 mm/day versus 0.48 mm/day for deterministic ACE2, reproduced precipitation power spectra and PDFs through the 99.99th percentile, and supplied coarse fields suitable for diffusion-based downscaling to 3 km precipitation (Perkins et al., 20 Dec 2025).
6. Scientific applications, methodological significance, and unresolved questions
Beyond emulation itself, ACE2 has been used as a computational instrument for otherwise prohibitive climate analyses. A prominent example is extreme-value estimation over the contiguous United States using a 10,560-year ACE2 ensemble trained on ERA5, generated from 12 initial states, 40 repetitions, and 22 years of recycled forcing data (Paciorek et al., 10 Oct 2025). In that study, ACE2 produced maxima that generally exceeded those in the ERA5 training data, enabling threshold-exceedance analyses of very rare precipitation and temperature events (Paciorek et al., 10 Oct 2025). The reported conclusions were that threshold-exceedance methods with sufficiently high thresholds are reliable for precipitation, that results are robust to season and storm-type variations, and that statistical uncertainty is well constrained, with relative standard error below 15% for most precipitation return values up to 1-in-100,000 years and below 5% for temperature even at million-year return values (Paciorek et al., 10 Oct 2025). The study also explicitly stated that it did not extensively investigate whether that specific emulator was fit for purpose, underscoring a broader methodological distinction between generating huge ensembles and validating their far-tail realism (Paciorek et al., 10 Oct 2025).
The ACE2 literature therefore occupies a dual role in climate science. On one side, it demonstrates that an autoregressive SFNO-based emulator can reproduce a wide range of atmospheric structures, trends, and responses at a throughput that makes century- to millennial-scale experimentation routine (Watt-Meyer et al., 2024). On the other, it has become a test case for the limits of historical-data training, forcing entanglement, and incomplete physical constraints. Across studies, recurrent unresolved issues include imperfect disentanglement of SST and CO4 effects, lack of explicit energy conservation in core ACE2, limited exposure to future-like climates, simplified or prescribed treatment of other Earth-system components, and inheritance of biases from the reference model or reanalysis used for training (Watt-Meyer et al., 2024, Landsberg et al., 26 Sep 2025, Clark et al., 2024, Clark et al., 6 Jun 2026).
Future development priorities are correspondingly specific rather than generic. The published literature calls for broader and more diverse training data, including non-equilibrium and intermediate forcing states; explicit enforcement of energy conservation; incorporation of additional Earth-system components such as interactive sea ice and more realistic ocean dynamics; exposure of more forcings than CO5; and evaluation suites that include process-based stress tests, regional trend benchmarks, and out-of-distribution experiments rather than mean-state metrics alone (Clark et al., 2024, Loon et al., 15 Feb 2025, Henn et al., 7 May 2026, Clark et al., 6 Jun 2026). ACE2 is thus best understood not simply as a fast climate model, but as a platform through which the scientific community is interrogating what atmospheric emulators can already do, what they still fail to do, and which physical and data-design choices most strongly determine that boundary.