- The paper demonstrates that end-conditioned cBottle-video ensembles can generate 1,000 atmospheric histories terminating at observed extremes while retaining 84–89% of the forward ensemble’s 500 hPa height spread.
- Reverse-time sampling reveals continuous alternative pathways that forward perturbation methods may miss, including persistent versus rapidly intensifying heatwave histories and Hurricane Ian tracks that bypass Florida before reaching the same final state.
- The method remains exploratory because generated trajectories lack guaranteed physical admissibility and calibrated probabilities, with underdispersion, wind–height inconsistencies, limited 66-hour context, and the need for physics-based validation.
Motivation and approach
Ensemble-based scenario planning for rare, high-impact weather events is limited by a combinatorial bottleneck: autoregressive emulators and physics-based models must be run forward from perturbed initial conditions, and the fraction of members that terminate in a specified extreme shrinks rapidly with lead time and rarity, forcing petabyte-scale ensembles to isolate a handful of relevant trajectories. Related approaches such as differentiable initial-condition optimization explore only a narrow neighborhood of a reference state rather than the full distribution of states consistent with an event. This paper inverts the problem: instead of asking what a given initial state produces, it asks which antecedent conditions could culminate in a specified extreme, by directly sampling the distribution of trajectories that terminate at a chosen state.
The enabling tool is cBottle-video, a score-based video diffusion model that generates twelve frames spanning a 66-hour sequence at 6-hour resolution in parallel, with arbitrary subsets of frames conditionable on existing data (2608.19008). The model uses an adapted Song-UNet diffusion backbone with learned position, time-of-day, and time-of-year embeddings, monthly AMIP SST conditioning, temporal self-attention, and a masked-conditioning training scheme that includes endpoint interpolation. It operates on a HEALPix (HPX64) grid at roughly 100 km resolution. The authors use a publicly available checkpoint zero-shot, subclassing earth2studio to expose the model's flexible conditioning, and generate 1000-member ensembles under three modes: end-conditioned (final frame pinned), start-conditioned (first frame pinned), and both-end conditioned (interpolation between two pinned states). Conditioning and verification use ERA5 reanalysis; tropical cyclone tracks are extracted with TempestExtremes.
Three case studies are analyzed: the 2021 Pacific Northwest (PNW) heatwave, Superstorm Sandy (2012), and Hurricane Ian (2022). A central quantitative finding is that end-conditioned ensembles retain 84–89% of the free-end 500 hPa geopotential height (z500​) spread of start-conditioned ensembles (0.84 for the heatwave, 0.86 for Ian, 0.89 for Sandy), demonstrating that reverse-time sampling does not collapse trajectory diversity.
The 2021 PNW heatwave
For the 66-hour window ending at the June 28, 2021 heat peak, end-conditioned members begin uniformly warmer than ERA5 and hold roughly steady rather than reproducing the observed rapid intensification; the warmest members exceed ERA5 by up to +5.8 K over land. The start-conditioned ensemble, in turn, underestimates the day-over-day rise in daily maxima and dampens the diurnal cycle. The generated anomalous warmth therefore manifests as an overly warm initial state with persistence, replacing the observed building heatwave with antecedent heat.
EOF analysis of the free-end z500​ field reveals a physically consistent but weak coupling: warmest-decile members feature a ridge anomaly over the Pacific Northwest, coolest-decile members a trough. However, the leading eight PCs explain only 9% of across-member variance in land-mean 2 m temperature (8% under 10-fold cross-validation), compressing a ~4 K spread into a ~1 K predicted range. The same eight modes capture 62% of the free-end height variance, so basis truncation is not the explanation. The authors attribute the weak coupling to the quasi-stationary blocking ridge: the heat dome's large-scale structure is shared across members and pinned by the end state, leaving only residual modulation to discriminate members. Notably, the start-conditioned regression is comparably weak (8% in-sample, 6% cross-validated), establishing near-symmetry in both directions. The authors concede that processes relevant to the residual spread—such as land feedbacks via soil moisture—are not represented in the model, since it lacks such variables.
Superstorm Sandy
For the 66-hour window spanning Sandy's approach and New Jersey landfall, end-conditioned members exhibit large antecedent diversity: storm centers begin a median of ~440 km from the observed starting position, with central pressures ranging from 958 to 998 hPa, and tracks form a smooth continuum with no discrete clustering. This is precisely the sampling that initial-condition perturbation cannot achieve; a 1000-member randomly perturbed ensemble of a Sandy-analog hurricane (Fiona) produced no Sandy-like outcomes in prior work, whereas gradient-based optimization explores only a narrow perturbation neighborhood.
The EOF regressions expose a marked conditioning-direction asymmetry. Under end-conditioning, the leading eight PCs explain 44% of antecedent storm latitude variance, 29% of distance from the observed fix, and ~15% each of longitude and intensity—and this organization is distributed across several modes with no single steering pattern, with some modes partly encoding storm position rather than ambient flow (meaning true environmental organization is weaker than reported). Under start-conditioning, the height field organizes the free end far more strongly: 78% of latitude variance, 63% of distance, 45% of longitude, and 20% of MSLP. EOF1 functions as a landfall-proximity axis along which ERA5 sits at the 0th percentile, and regression evaluated at ERA5's synoptic state extrapolates to a predicted landfall distance of ~0 km—yet no member samples that extreme tail. This null result is consistent with the ~700-year return period of Sandy's track and the difficulty of sampling it with forward ensembles. No start-conditioned member reproduces Sandy's rapid intensification: none is as deep as the observed 950.7 hPa.
Hurricane Ian
For Ian's window spanning both U.S. landfalls, the end-conditioned ensemble shows a consequential bifurcation: 88.4% of members make an intermediate Florida landfall, while 10.9% bypass Florida entirely and travel north along the Atlantic coast before reaching the pinned South Carolina state. Because all analyzable members are anchored to the observed final landfall, bypass tracks represent counterfactual trajectories that reach Ian's terminal state without the Florida impacts—underscoring that intermediate hazard exposure varies substantially across plausible histories of the same terminal event.
The start-versus-end asymmetry is sharper than for Sandy. End-conditioned EOF regressions explain no more than ~5% of antecedent latitude, longitude, and distance variance (13% for intensity), despite the leading eight modes capturing 73% of the height field variance. Under start-conditioning, the same regressions explain 48% of free-end longitude, 33% of latitude, and 21% of distance. Two leading modes separate cleanly: EOF1 (18.4% of height variance) sorts members along the coast, and EOF2 (13.3%) sorts them shoreward, with ERA5's height state at the 96.6th percentile of PC2—the same edge geometry seen for Sandy. Ian's intensity exceeds 98% of members in both conditioning directions, again placing the realized event in the ensemble tail.
Physical admissibility and verification
The paper is explicit that generated trajectories cannot be assumed physically admissible. Because the emulator optimizes pixel-level score matching without enforcing dynamical conservation laws, it will construct a transition between any pair of conditioned states. A deliberate negative control makes this concrete: conditioning the window on two tropical cyclones from different years on opposite sides of the equator, the model smoothly dissolves one vortex and nucleates the other rather than signaling inadmissibility. In the case studies the conditioning states were dynamically contiguous, but the absence of an intrinsic admissibility filter is a structural limitation of the approach.
Diagnostic checks reveal systematic deficiencies of the off-the-shelf checkpoint. Across all cases, midlatitude 500 hPa wind magnitude weakens away from the pinned end, accompanied by a rising ageostrophic fraction—for Ian and the heatwave, an ageostrophic component opposing the geostrophic flow leaves the wind more subgeostrophic than the height field implies, surfacing a physical inconsistency between z500​ and the winds. The 500 hPa spread-to-skill ratio never exceeds 0.9 in the synoptic domains, indicating structural underdispersion and ensemble-mean bias, although the spread-to-error ratio for land 2 m temperature does reach or exceed unity, showing that underdispersion is not universal across variables. The authors frame these as limitations of zero-shot evaluation rather than intrinsic failures of end-conditioning, and note candidate remedies: physics-constrained flow matching, and bidirectional generative models that can self-supervise rollout error. They further argue that final validation likely requires forward-integrating the discovered initial conditions in a physics-based model to confirm convergence toward the observed end state.
Limitations and open questions
Several constraints bound the results. The 66-hour window is imposed by the cBottle-video context window; extending histories further back via autoregressive conditioning is untested, and conditioned frames are reproduced with subtle but noticeable degradation, so compounding errors in longer rewinds remain an open problem. The checkpoint was not trained for or evaluated on autoregressive performance. The EOF regressions measure organization rather than steering causality, since predictors and targets share the same frame; storm-position leakage in some modes means reported environmental organization is an upper bound. Tracker-based exclusions (up to 4.4% of end-conditioned Sandy members, attributable to warm-core loss during extratropical transition) and nearest-valid-fix substitutions for members without an exact-frame fix introduce additional measurement uncertainty. The verification diagnostics cover one level, one band, and one balance relation, and the authors call for a more comprehensive suite of dynamical, spectral, and conservation metrics. Finally, the method as presented yields plausible histories without calibrated probabilities; pairing arbitrary-boundary sampling with guided likelihood estimation is proposed but not demonstrated here.
Conclusion
This work demonstrates direct generative sampling of the distribution of atmospheric histories terminating at a realized extreme, using end-conditioning in a video diffusion emulator applied to three historical disasters. End-conditioned 1000-member ensembles retain 84–89% of forward-ensemble synoptic spread, and the resulting trajectories reveal materially distinct pathways—persistent versus rapidly intensifying heat for the PNW event, divergent intermediate landfalls for Ian—while the realized events sit in the tails of both conditioning directions. The large-scale height field organizes only a minority of the antecedent diversity (5–44%), with alternative histories forming a continuous spectrum rather than discrete synoptic clusters, a structure that regime catalogs cannot enumerate and that only direct sampling can access. Whether generated rewinds are dynamically admissible, and how to attach calibrated probabilities to them, remain the principal unresolved questions before the method can support operational climate risk assessment.