---
title: 'CES: A Bayesian Inversion Framework'
url: https://www.emergentmind.com/topics/calibrate-emulate-sample-ces
type: topic
---

# CES: A Bayesian Inversion Framework

Searching arXiv for CES and related calibrate-emulate-sample papers.
Calibrate–Emulate–Sample (CES) is a three-stage framework for Bayesian inversion and Bayesian calibration in settings where the forward or parameter-to-data map is expensive to evaluate, derivatives or adjoints are unavailable, and forward evaluations may be noisy. In its canonical formulation, CES first uses an ensemble-based inverse method to move parameter ensembles toward regions of high posterior probability, then fits a surrogate model—typically a Gaussian process (GP)—to the resulting simulator evaluations, and finally samples an approximate Bayesian posterior with MCMC using the surrogate in place of the original simulator [2001.03689]. The framework was introduced as a derivative-free strategy for approximate Bayesian inference requiring only a small number of forward evaluations and has since been used in climate-model calibration and experimental design, where it enables uncertainty quantification with approximately $\mathcal{O}(10^2)$ simulator evaluations rather than the $\mathcal{O}(10^5)$ evaluations often associated with direct MCMC on the full model [2001.03689], [2012.13262], [2201.06998].

## 1. Bayesian inverse-problem setting

CES is formulated for inverse problems of the form
$$
y = G(\theta) + \eta,
$$
where $\theta \in \mathbb{R}^p$ denotes unknown parameters, $y \in \mathbb{R}^d$ observed data, $G:\mathbb{R}^p \to \mathbb{R}^d$ the forward or parameter-to-data map, and $\eta$ observational noise [2001.03689]. In the simplest Bayesian setting described in the original CES formulation, both the noise and the prior are Gaussian, leading to a posterior density
$$
\pi^y(\theta) \propto \exp\bigl(-\Phi_R(\theta)\bigr),
$$
with negative log-posterior
$$
\Phi_R(\theta)=\frac{1}{2}\|y-G(\theta)\|^2_{y}+\frac{1}{2}\|\theta\|^2.
$$
The norm is defined by $\|v\|_A := \|A^{-1/2}v\|$, so the data misfit is weighted by the observation covariance [2001.03689].

The motivating regime is one in which evaluating $G(\theta)$ is very expensive, gradients or derivative-adjoints are unavailable, and in some applications only noisy evaluations of $G$ are accessible [2001.03689]. Standard MCMC remains the reference mechanism for principled posterior sampling and uncertainty quantification, but its evaluation cost can be prohibitive for large PDE and climate models. CES addresses this by separating the inferential task into a low-budget calibration stage, an emulator construction stage, and a posterior sampling stage carried out on the emulator rather than on the original simulator [2001.03689], [2012.13262].

In climate applications, the same logic is expressed in terms of time-averaged statistics. For finite averaging window $T$, the forward map is written as $\mathcal{G}_T(\theta;z^{(0)})$, while the target statistical map is $\mathcal{G}_\infty(\theta)$; internal variability is modeled through a central-limit-theorem approximation
$$
\mathcal{G}_T(\theta;z^{(0)}) \approx \mathcal{G}_\infty(\theta) + \sigma,\qquad \sigma\sim N(0,\Sigma),
$$
which makes direct likelihood-based MCMC both noisy and expensive [2012.13262]. This use case exemplifies the class of problems for which CES was designed.

## 2. Three-stage architecture

CES consists of three coupled stages: calibration, emulation, and sampling [2001.03689]. The stages are conceptually distinct but computationally interdependent, because the simulator evaluations generated during calibration are reused as training data for the emulator, and the calibrated ensemble also supplies an initial state and scale information for MCMC in the sampling stage [2001.03689].

The original methodological statement is that CES first uses an ensemble Kalman method to calibrate parameters and locate the high-posterior region cheaply, second uses those forward evaluations to emulate the map with a surrogate model, and third uses the surrogate inside MCMC to sample an approximate Bayesian posterior at low cost [2001.03689]. In climate-model applications, this same structure was implemented with ensemble Kalman inversion (EKI) for calibration, GP emulation of time-averaged model outputs, and MCMC on the emulator to approximate the posterior over uncertain convective parameters [2012.13262]. In the subsequent experimental-design work, CES became the computational engine for repeated posterior estimation over many candidate design locations and seasons [2201.06998].

The practical significance of this decomposition is that the expensive simulator is used only in the early phase, often on the order of $\mathcal{O}(10^2)$ evaluations, whereas the posterior-sampling phase is shifted to a cheap surrogate [2001.03689], [2201.06998]. A central claim in the original paper is that ensemble Kalman sampling (EKS) solves, cheaply and derivative-free, the design problem of where to place points in parameter space to train the emulator efficiently for Bayesian inversion [2001.03689]. This distinguishes CES from workflows based solely on space-filling designs or from methods that attempt direct MCMC on the full simulator.

A recurrent misconception is that CES is merely a surrogate-modeling pipeline. In the cited formulations, the method is more specific: the calibration stage is not only a preliminary optimizer but also an adaptive design mechanism targeted at the main support of the posterior, and the sampling stage is intended to recover approximate Bayesian uncertainty quantification rather than only point estimates [2001.03689], [2012.13262].

## 3. Calibration stage: ensemble-based localization of posterior mass

In the original formulation, the calibration stage uses ensemble Kalman sampling (EKS), a stochastic interacting-particle dynamics for ensemble members $\{\theta^{(j)}\}_{j=1}^J$:
$$
\frac{d\theta^{(j)}}{dt} = -\frac{1}{J}\sum_{k=1}^J \left\langle G(\theta^{(k)})-\bar G,\; G(\theta^{(j)})-y\right\rangle_{y} \left(\theta^{(k)}-\bar\theta\right) - C(\Theta)^{-1}\theta^{(j)} + \sqrt{2C(\Theta)}\,\frac{dW^{(j)}}{dt},
$$
with
$$
\bar\theta=\frac{1}{J}\sum_{k=1}^J \theta^{(k)}, \qquad \bar G=\frac{1}{J}\sum_{k=1}^J G(\theta^{(k)}),
$$
and
$$
C(\Theta)=\frac{1}{J}\sum_{k=1}^J (\theta^{(k)}-\bar\theta)\otimes(\theta^{(k)}-\bar\theta).
$$
The first term drives data fit, the $-C(\Theta)^{-1}\theta^{(j)}$ term incorporates the prior, and the Brownian term makes the ensemble behave like a sampler rather than a pure optimizer [2001.03689].

The paper discretizes this with a linearly implicit split-step scheme:
$$
{\theta}^{(*, j)}_{n+1} = {\theta}^{(j)}_{n} -\Delta t_n \frac{1}{J}\sum_{k=1}^J \left\langle G(\theta^{(k)}_n)-\bar G,\; G(\theta^{(j)}_n)-y\right\rangle_{y} \,\theta^{(k)}_n -\Delta t_n C(\Theta_n)^{-1}{\theta}^{(*, j)}_{n+1},
$$
$$
{\theta}^{(j)}_{n+1} = {\theta}^{(*, j)}_{n+1} +\sqrt{2\Delta t_n\,C(\Theta_n)}\,\xi_n^{(j)}, \qquad \xi_n^{(j)}\sim N(0,I).
$$
This stage is derivative-free and uses only forward evaluations of the simulator [2001.03689].

In the climate-model variant, the calibration stage is implemented instead with EKI. The finite-time data model is
$$
y = \mathcal{G}_T(\theta;z^{(0)}) + \eta,\qquad \eta\sim N(0,\Delta),
$$
with calibration misfit
$$
\Phi_T(\theta, y;w^{(0)}) = \frac{1}{2}\| y - \mathcal{G}_T(\theta;w^{(0)})\|_{\Gamma+\Sigma}^2.
$$
The EKI update is
$$
\theta_{m}^{(n+1)} = \theta_{m}^{(n)} + C_{\theta \mathcal{G}^{(n)} \left( \Gamma + C_{\mathcal{G}\mathcal{G}^{(n)} \right)^{-1} \left( y - \mathcal{G}_T(\theta_{m}^{(n)}) \right),
$$
and is described as derivative-free, ensemble-based, robust to noisy forward evaluations, and computationally cheap relative to brute-force Bayesian sampling [2012.13262].

The principal role of the calibration stage across these CES variants is to concentrate parameter ensembles in regions of high posterior probability or near the best-fitting parameter region [2001.03689], [2201.06998]. That property is methodologically crucial because the emulator is subsequently trained where approximation quality matters most for inference. At the same time, the climate-model paper explicitly notes an important limitation: EKI is an optimization method and its collapsing ensemble spread is not a reliable posterior uncertainty estimate [2012.13262]. CES therefore does not equate ensemble spread in the calibration phase with posterior uncertainty; instead, it reserves uncertainty quantification for the final sampling phase.

## 4. Emulation stage: surrogate construction in posterior-relevant regions

After calibration, CES constructs a surrogate for the parameter-to-data map from the expensive forward evaluations already obtained. In the original paper, the calibration output provides training data
$$
\{\theta^{(m)},G(\theta^{(m)})\}_{m=1}^M,
$$
with $M \le JN$ in general, and often $M=J$ by using the final EKS ensemble as the design set [2001.03689]. The stated rationale is that the EKS output is concentrated in regions of high posterior mass, so the emulator is trained precisely where accurate approximation matters most for Bayesian inference [2001.03689].

The default emulator in CES is a Gaussian process, often fitted independently to each output component $G_l(\theta)$, $l=1,\dots,d$, using a squared-exponential kernel with a nugget:
$$
k_l(\theta,\theta') = \sigma_l^2 \exp\!\left(-\frac12\|\theta-\theta'\|_{D_l}^2\right) +\lambda_l^2\,\delta_\theta(\theta').
$$
Here $\sigma_l^2$ is the amplitude, $D_l$ encodes lengthscales, and $\lambda_l^2$ represents observation or evaluation noise in the simulator outputs [2001.03689]. The emulator is represented as
$$
G^{(M)}(\theta)\sim N\bigl(m(\theta),\Gamma(\theta)\bigr),
$$
so the forward map is replaced by a probabilistic surrogate with predictive mean $m(\theta)$ and covariance $\Gamma(\theta)$ [2001.03689].

The original CES paper assigns two functions to the emulator: cheap evaluation inside MCMC and noise reduction or smoothing when the original simulator is noisy, as in finite-time averages in chaotic systems [2001.03689]. In climate CES, the emulator is trained on finite-time outputs $\mathcal{G}_T(\theta)$ but is intended to approximate the infinite-time or statistical map $\mathcal{G}_\infty(\theta)$ [2012.13262], [2201.06998]. This shift from noisy finite-time evaluations to a smooth GP surrogate is one reason the final MCMC stage is more stable than direct sampling on the simulator.

Several papers emphasize decorrelation transforms before GP fitting. The original CES paper discusses diagonalizing the observation covariance or using PCA/SVD on the EKS-generated training outputs so that independent GP emulators can be trained in transformed coordinates [2001.03689]. The idealized GCM implementation diagonalizes $\Sigma$ through
$$
\Sigma = V D^2 V^T,\qquad \tilde{Y} = D^{-1}V^T Y,
$$
and then trains scalar-valued GPs on the decorrelated components [2012.13262]. The experimental-design paper likewise uses an SVD of the internal-variability covariance,
$$
\Sigma \approx V^\top \Lambda V,
$$
and builds the emulator in a transformed coordinate system [2201.06998].

A plausible implication is that CES is not tied to a particular GP architecture so much as to a specific local-design principle: surrogate accuracy is prioritized in the posterior-relevant region rather than globally over the prior domain. That interpretation is explicit in the claim that EKS provides a cheap solution to the experimental-design problem of emulator training [2001.03689].

## 5. Sampling stage: approximate posterior inference on the surrogate

Once the forward map is replaced by the emulator, CES samples from an approximate posterior in which $G$ is replaced by $G^{(M)}$:
$$
y = G^{(M)}(\theta) + \eta.
$$
Depending on the treatment of emulator and observation uncertainty, the approximate negative log-likelihood takes several forms [2001.03689]. If emulator uncertainty is ignored,
$$
\Phi^{(M)}(\theta)=\frac12\|y-m(\theta)\|^2_{y}.
$$
If simulator uncertainty is retained but observation noise is neglected,
$$
\Phi^{(M)}(\theta)=\frac12\|y-m(\theta)\|^2_{\Gamma(\theta)}+\frac12\log\det\Gamma(\theta).
$$
If both are included,
$$
\Phi^{(M)}(\theta)= \frac12\|y-m(\theta)\|^2_{\Gamma(\theta)+y} +\frac12\log\det\bigl(\Gamma(\theta)+y\bigr).
$$
The resulting approximate posterior is
$$
\pi^{(M)}(\theta)\propto \exp\!\bigl(-\Phi^{(M)}(\theta)\bigr).
$$
MCMC is then run on this surrogate posterior rather than on the original model [2001.03689].

In the original CES implementation, a random-walk Metropolis method is initialized at the final EKS ensemble mean,
$$
\theta_0=\bar\theta_J,
$$
with Gaussian proposal covariance taken from the final EKS ensemble covariance:
$$
\theta^*_{n+1}=\theta_n+\xi_n,\qquad \xi_n\sim N(0,C(\Theta_J)).
$$
The acceptance probability is
$$
a(\theta,\theta^*)= \min\left\{ 1,\exp\left[ \bigl(\Phi^{(M)}(\theta^*)+\tfrac12\|\theta^*\|^2\bigr) - \bigl(\Phi^{(M)}(\theta)+\tfrac12\|\theta\|^2\bigr) \right] \right\}.
$$
Thus the calibration stage informs not only emulator training but also MCMC initialization and proposal scale [2001.03689].

In the climate-model formulation, the posterior in decorrelated space is written as
$$
\mathbb{P}(\theta \mid \tilde{y}) \propto \mathbb{P}(\tilde{y}\mid\theta)\mathbb{P}(\theta),
$$
with emulator-based likelihood
$$
\mathbb{P}(\theta \mid \tilde{y}) \propto \exp\left(-\frac{1}{2}\| \tilde{y} - \tilde{\mathcal{G}_{\mathrm{GP}(\theta)}\|^2_{\tilde{\Gamma}_{\mathrm{GP}(\theta)} - \frac{1}{2}\log \det \tilde{\Gamma}_{\text{GP}(\theta)} \right)\mathbb{P}(\theta),
$$
and, for Gaussian prior $N(m,C)$,
$$
\Phi_{\mathrm{MCMC}(\theta,\tilde{y}) = \frac{1}{2}\| \tilde{y} - \tilde{\mathcal{G}_{\mathrm{GP}(\theta)}\|^2_{\tilde{\Gamma}_{\mathrm{GP}(\theta)} + \frac{1}{2}\log \det \tilde{\Gamma}_{\text{GP}(\theta)} + \frac{1}{2}\|\theta-m\|^2_C.
$$
The log-determinant term is retained because the covariance depends on $\theta$ [2012.13262]. This is a more explicit treatment of emulator uncertainty than is often used in simpler surrogate-based workflows.

A common misunderstanding is that the sampling stage merely produces draws from the GP emulator. In CES, the target is an approximate Bayesian posterior over the parameters, and the emulator serves only as a substitute for the expensive forward map inside the likelihood [2001.03689], [2012.13262].

## 6. Applications, numerical studies, and extensions

The original CES paper validates the framework on a linear inverse problem, a Darcy flow inverse problem, and Lorenz ’63 and Lorenz ’96 systems [2001.03689]. In the linear inverse problem, the exact posterior is known, EKS calibrates correctly, GP emulation becomes accurate as the training set grows, and GP-based MCMC recovers the posterior well using $M=J$ design points [2001.03689]. In the Darcy flow inverse problem, CES with $J=512$ closely matches gold-standard MCMC, while with $J=128$ some systematic deviation appears although posterior mass still captures the truth [2001.03689]. In Lorenz ’63 and Lorenz ’96, time-averaged observations generate noisy forward evaluations; the GP stage smooths the misfit and improves sampling, and for Lorenz ’96 the method is particularly useful because gold-standard MCMC is too expensive or rejects too often [2001.03689].

In climate science, CES was developed into a practical framework for calibrating an idealized GCM with internal variability noise [2012.13262]. The uncertain parameters were the convective relaxation timescale $\tau$ and the relative humidity parameter RH, with priors
$$
\theta_{\mathrm{RH} \sim \mathrm{Logitnormal}(0, 1), \qquad \theta_{\tau} \sim \mathrm{Lognormal}(12~\mathrm{h}, (12~\mathrm{h})^2),
$$
implemented through unconstrained transformed coordinates [2012.13262]. Synthetic data were generated in a perfect-model setting from 30-day zonal means of free-tropospheric relative humidity, precipitation rate, and extreme precipitation frequency, all at 32 latitudes [2012.13262]. The paper reports that the CES approach generates parameter distributions that approximate the Bayesian posteriors at a fraction of the usual computational cost, with about 600 short GCM runs and an estimated factor 1000 speedup relative to direct sampling on the GCM [2012.13262].

CES was subsequently embedded in an ensemble-based experimental-design algorithm for climate-model parameterizations [2201.06998]. In that setting, CES is run repeatedly over candidate design points $k$ to estimate the posterior
$$
\pi(\boldsymbol{\theta}\mid W_k\mathbf{y}),
$$
compute the utility
$$
U(W_k) = \left( \det\big(\mathrm{Cov}(\boldsymbol{\theta}\mid W_k\mathbf{y})\big) \right)^{-1},
$$
and select the maximizing design
$$
\tilde{k}=\arg\max_{k\in D} U(W_k).
$$
The paper interprets this as information gain through reduction in posterior uncertainty volume and reports that the largest information gain typically, but not always, results from regions near the intertropical convergence zone [2201.06998]. This work extends CES from a calibration methodology to an engine for Bayesian optimal data targeting.

A broader extension is suggested by a 2025 comparison study of surrogate-based Bayesian calibration methods on the Lorenz ’96 multiscale system [2508.13071]. There CES is described as a sequential Bayesian calibration framework using EKS, GP emulation, and MCMC, and is compared with History Matching (HM), Bayesian Optimal Experimental Design (BOED), and Goal-Oriented BOED (GBOED). The study finds that CES offers excellent performance but at high computational expense, while GBOED achieves comparable accuracy using fewer model evaluations [2508.13071]. This suggests that later work increasingly treats CES as a baseline architecture against which more explicitly design-optimized strategies are compared.

## 7. Advantages, limitations, and related interpretations

The main advantages attributed to CES in the original and subsequent papers are consistent. It is derivative-free, requires only a small number of expensive forward solves relative to direct MCMC, provides a good design for emulator training because the calibration stage targets high-posterior regions, supports approximate Bayesian uncertainty quantification through posterior sampling, and is robust to noisy forward models because the GP emulator smooths simulator noise [2001.03689]. In climate applications, CES is also presented as robust to internal climate variability and capable of converting parameter uncertainty into predictive uncertainty through posterior propagation [2012.13262].

Its limitations are equally explicit. The calibration ensemble spread from EKI should not be interpreted as posterior uncertainty, because EKI is fundamentally an optimization method whose ensemble collapses toward consensus [2012.13262]. The GP emulator can become the main scalability bottleneck in higher-dimensional parameter spaces [2012.13262]. The method works best when the EKS or EKI stage succeeds in locating the region that carries most posterior mass; if calibration misses relevant regions, the local emulator may be inadequate there [2001.03689], [2012.13262]. The 2025 comparison study further indicates that CES can be computationally expensive, and on the Lorenz ’96 multiscale benchmark EKS alone can outperform CES for that specific testbed [2508.13071].

Several adjacent uses of the acronym “CES” are not the same method. The 2017 paper on bilateral multifactor CES general equilibrium concerns constant elasticity of substitution calibration and uses “CES” in the sense of CES aggregators, not Calibrate–Emulate–Sample [1706.09365]. By contrast, the 2023 paper on Bayesian neural networks adopts a “Calibration-Emulation-Sampling” strategy for posterior approximation in BNNs, but its calibration stage uses early-stopped SGHMC, its emulator is a neural network approximating the parameter-to-output or likelihood map, and its sampling stage uses pCN rather than the original EKS–GP–MCMC structure [2312.11799]. This suggests that “CES” has broadened from a specific ensemble-Kalman-based surrogate pipeline into a more general design pattern for expensive Bayesian inference, though the canonical meaning remains the three-stage framework introduced for Bayesian inversion in 2020 [2001.03689].

A plausible implication of the later literature is that CES occupies a middle position between purely optimization-based calibration and fully design-theoretic experimental design. It retains posterior sampling and uncertainty quantification, unlike pure calibration or history-matching workflows, but it avoids the cost of direct MCMC on the simulator by concentrating computational effort into a posterior-relevant surrogate. That balance explains its continuing use as both a practical inference method and a reference point for more specialized calibration and experimental-design schemes [2001.03689], [2201.06998], [2508.13071].

Source: https://www.emergentmind.com/topics/calibrate-emulate-sample-ces