---
title: Controlled Sequential Monte Carlo
url: https://www.emergentmind.com/topics/controlled-sequential-monte-carlo
type: topic
---

# Controlled Sequential Monte Carlo

Controlled Sequential Monte Carlo (cSMC) is a Sequential Monte Carlo methodology in which the proposal mechanism is improved by solving, approximately and iteratively, an associated optimal control / dynamic programming problem [1708.08396]. In the modern formulation, cSMC replaces a fixed proposal sequence by a sequence of twisted path measures parameterized by a policy, and seeks a policy that makes the SMC weights as flat as possible—ideally constant, yielding a zero-variance normalizing-constant estimator [1708.08396]. The term is also used more loosely for neighboring methods that modify proposal kernels, mutation kernels, resampling laws, or intermediate targets to reduce mismatch between particle trajectories and the target distribution, but several of these methods are not control-theoretic cSMC in the strict sense [1506.03338][0808.3466][1607.05758].

## 1. Formal setting and scope

Controlled Sequential Monte Carlo is most naturally stated in the Feynman–Kac framework. Let \(X_{0:T}=(X_0,\dots,X_T)\) be a non-homogeneous Markov chain on \((\mathsf X,\mathcal X)\) with law
\[
\mathbb Q(\mathrm d x_{0:T}) = \mu(\mathrm d x_0)\prod_{t=1}^T M_t(x_{t-1},\mathrm d x_t),
\]
and let strictly positive potentials \(G_0\in\mathcal B(\mathsf X)\) and \(G_t\in\mathcal B(\mathsf X\times\mathsf X)\), \(t\ge 1\), define the path measure
\[
\mathbb P(\mathrm d x_{0:T}) = Z^{-1} G_0(x_0)\prod_{t=1}^T G_t(x_{t-1},x_t)\,\mathbb Q(\mathrm d x_{0:T}),
\]
with normalizing constant
\[
Z=\mathbb E_{\mathbb Q}\!\left[G_0(X_0)\prod_{t=1}^T G_t(X_{t-1},X_t)\right].
\]
The associated unnormalized and normalized marginals are
\[
\gamma_t(\varphi)=\mathbb E_{\mathbb Q}\!\left[\varphi(X_t)G_0(X_0)\prod_{s=1}^t G_s(X_{s-1},X_s)\right],\qquad
\eta_t(\varphi)=\frac{\gamma_t(\varphi)}{Z_t},
\]
with \(Z_t=\gamma_t(\mathsf X)\) and final target \(\eta_T\) [1708.08396].

This formulation covers both state-space and static models. In state-space models, \(\mathbb P\) is the smoothing distribution and \(Z\) the marginal likelihood. In static models, \(\eta_T\) is the target posterior or target distribution, obtained via a sequence of bridging measures [1708.08396]. Tutorial treatments place cSMC within the broader auxiliary or twisted SMC methods literature and emphasize that proposal distributions and intermediate target distributions are the two main user design choices of SMC [1903.04797]. The general SMC-sampler substrate is the sequence \((\pi_t)\), forward kernels \((M_t)\), backward kernels \((L_t)\), and incremental weight
\[
w_t(x_{t-1},x_t)=\frac{\gamma_t(x_t)L_{t-1}(x_t,x_{t-1})}{\gamma_{t-1}(x_{t-1})M_t(x_{t-1},x_t)},
\]
which is the basic mechanism from which more elaborate twisted and controlled constructions are built [2007.11936].

## 2. Twisting, optimal control, and the zero-variance benchmark

The distinctive move in cSMC is to introduce a policy \(\psi=(\psi_t)_{t=0}^T\) of positive functions and to twist the proposal path measure. Given any path measure
\[
\mathbb F(\mathrm d x_{0:T})=\nu(\mathrm d x_0)\prod_{t=1}^T K_t(x_{t-1},\mathrm d x_t),
\]
its \(\psi\)-twisted version is
\[
\mathbb F^\psi(\mathrm d x_{0:T})=\nu^\psi(\mathrm d x_0)\prod_{t=1}^T K_t^\psi(x_{t-1},\mathrm d x_t),
\]
where
\[
\nu^\psi(\mathrm d x_0)=\frac{\nu(\mathrm d x_0)\psi_0(x_0)}{\nu(\psi_0)},\qquad
K_t^\psi(x_{t-1},\mathrm d x_t)=\frac{K_t(x_{t-1},\mathrm d x_t)\psi_t(x_{t-1},x_t)}{K_t(\psi_t)(x_{t-1})}.
\]
Applied to \(\mathbb Q\), this yields a controlled proposal path measure \(\mathbb Q^\psi\) [1708.08396].

Under \(\mathbb Q^\psi\), the same target \(\mathbb P\) can be rewritten with twisted potentials:
\[
G_0^\psi(x_0)=\frac{\mu(\psi_0)G_0(x_0)M_1(\psi_1)(x_0)}{\psi_0(x_0)},
\]
\[
G_t^\psi(x_{t-1},x_t)=\frac{G_t(x_{t-1},x_t)M_{t+1}(\psi_{t+1})(x_t)}{\psi_t(x_{t-1},x_t)},\qquad t=1,\dots,T-1,
\]
\[
G_T^\psi(x_{T-1},x_T)=\frac{G_T(x_{T-1},x_T)}{\psi_T(x_{T-1},x_T)}.
\]
The terminal normalizing constant is preserved:
\[
Z=Z_T^\psi.
\]
Hence cSMC changes the simulation law without changing the target [1708.08396].

The optimal policy is characterized by a backward recursion. If \(\psi\) is a current policy and \(\phi\) is an additional twist, the optimal \(\phi^*\) solves
\[
\inf_{\phi\in\Phi}\mathrm{KL}\!\left((\mathbb Q^\psi)^\phi \mid \mathbb P\right),
\]
with recursion
\[
\phi_T^*(x_{T-1},x_T)=G_T^\psi(x_{T-1},x_T),
\]
\[
\phi_t^*(x_{t-1},x_t)=G_t^\psi(x_{t-1},x_t)\,M_{t+1}^\psi(\phi_{t+1}^*)(x_t),\qquad t=1,\dots,T-1,
\]
\[
\phi_0^*(x_0)=G_0^\psi(x_0)\,M_1^\psi(\phi_1^*)(x_0).
\]
Equivalently, the optimal value functions \(V_t^*=-\log\phi_t^*\) satisfy a Bellman recursion [1708.08396].

This recursion yields the zero-variance benchmark. If \(\psi^*=\psi\cdot\phi^*\), then \(\mathbb Q^{\psi^*}=\mathbb P\), the twisted marginals coincide with the true marginals of \(\mathbb P\), and
\[
Z^{\psi^*,N}=Z \quad \text{almost surely for every } N.
\]
Substituting \(\phi^*\) into the twisted potentials gives
\[
G_0^{\psi^*}(x_0)=Z,\qquad G_t^{\psi^*}(x_{t-1},x_t)=1,\quad t=1,\dots,T,
\]
so all weights are constant [1708.08396]. In state-space language, this is the controlled analogue of a fully adapted particle system; in general Feynman–Kac language, it is the ideal twisted path measure.

## 3. Approximate dynamic programming, policy classes, and online extensions

Exact optimal policies are generally intractable, so practical cSMC replaces the Bellman recursion by approximate dynamic programming. The basic algorithm starts from
\[
\psi^{(0)}\equiv 1,
\]
which corresponds to uncontrolled SMC, runs a \(\psi^{(i)}\)-twisted SMC pass, and then regresses the value functions \(V_t^*=-\log \phi_t^*\) on a tractable function class \(\mathsf F_t\) [1708.08396]. At time \(T\), one solves
\[
\hat V_T=\arg\min_{\varphi\in\mathsf F_T}\sum_{n=1}^N\left(\varphi(X_{T-1}^{A_{T-1}^n},X_T^n)+\log G_T^\psi(X_{T-1}^{A_{T-1}^n},X_T^n)\right)^2,
\]
and backward in time one fits
\[
\hat V_t=\arg\min_{\varphi\in\mathsf F_t}\sum_{n=1}^N\left(\varphi(X_{t-1}^{A_{t-1}^n},X_t^n)+\log \xi_t(X_{t-1}^{A_{t-1}^n},X_t^n)\right)^2,
\]
where
\[
\xi_t(x_{t-1},x_t)=G_t^\psi(x_{t-1},x_t)\,M_{t+1}^\psi(\hat\phi_{t+1})(x_t).
\]
The policy is refined multiplicatively,
\[
\psi^{(i+1)}=\psi^{(i)}\cdot \hat\phi^{(i+1)},
\]
and the residuals determine how close the new twisted potentials are to constants [1708.08396].

This ADP view is central: cSMC is approximate dynamic programming for a finite-horizon KL control problem on path space [1708.08396]. Function classes include quadratic functions in state-space models, quadratic-plus-likelihood terms for static models, radial basis functions for multimodal policies, and linear least-squares basis expansions [1708.08396]. The analysis introduces Bellman semigroups, residual-based KL discrepancy bounds, semigroup stability constants, and a CLT for the estimated policy under linear least-squares classes [1708.08396].

Application-specific implementations exploit tractable policy classes. In diffusion-process inference via splitting schemes, the controlled proposal is built on conditionally Gaussian kernels and quadratic log-policies
\[
\phi_k(x_k)=\log\psi_k(x_k)=x_k^\top A_k x_k + x_k^\top b_k + c_k,
\]
with
\[
\mathcal F_k^{\mathrm{eq}}=\left\{\phi_k(x_k)=x_k^\top A_k x_k+x_k^\top b_k+c_k : (A_k,b_k,c_k)\in \mathbb S_d\times\mathbb R^d\times\mathbb R \right\}.
\]
Because multiplying a Gaussian kernel by \(\exp(\phi_k)\) yields another Gaussian kernel, \(M_k^\psi\) is tractable and \(M_k(\psi_k)\) can be computed analytically [2507.14535]. In that setting, cSMC is applied to Feynman–Kac representations of pseudolikelihoods arising from semi-linear SDEs, partial observation, and bridge augmentation.

Recent online work extends the offline smoothing-oriented formulation to real-time hidden Markov models through Online Rolling Controlled Sequential Monte Carlo (ORCSMC) [2508.00696]. ORCSMC defines a rolling window of length \(L\),
\[
t_0\coloneqq \max\{1,t-L+1\},
\]
and uses two particle systems: a learning filter that estimates twisting functions over the current window and an estimation filter that delivers filtering, smoothing, and marginal-likelihood output [2508.00696]. The twisted model is expressed through positive functions \(\psi_t:\mathbb R^d\to(0,\infty)\),
\[
\mu^\psi(x_1)=\frac{\mu(x_1)\psi_1(x_1)}{\mu(\psi_1)},\qquad
f_t^\psi(x_t\mid x_{t-1})=\frac{f_t(x_t\mid x_{t-1})\psi_t(x_t)}{f_t(\psi_t)(x_{t-1})},
\]
with corresponding twisted potentials
\[
g_1^\psi(x_1)=\frac{g_1(y_1\mid x_1)\mu(\psi_1)}{\psi_1(x_1)}\,f_2(\psi_2)(x_1),\qquad
g_t^\psi(x_t)=\frac{g_t(y_t\mid x_t)}{\psi_t(x_t)}\,f_{t+1}(\psi_{t+1})(x_t),
\]
and with local backward recursion targets
\[
\tilde\psi_t^n \leftarrow g_t(y_t\mid X_t^n)\, f_{t+1}(\psi_{t+1})(X_t^n).
\]
For fixed lag \(L\), particle count \(N\), and iteration count \(K\), the method enforces bounded per-observation computation and memory [2508.00696].

## 4. Neighboring methodologies and historical lineage

The strict control-theoretic meaning of cSMC should be distinguished from several adjacent literatures that pursue related goals. Neural Adaptive Sequential Monte Carlo (NASMC) learns a forward proposal family by minimizing the inclusive KL divergence
\[
\mathrm{KL}[p_\theta(z_{1:T}\mid x_{1:T}) \,\|\, q_\phi(z_{1:T}\mid x_{1:T})],
\]
with particle-weighted gradients of the form
\[
\sum_t \sum_n \tilde{w}_t^{(n)} \nabla_\phi \log q_\phi(z_t^{(n)}\mid \cdot).
\]
It is aligned in goal with controlled SMC because it modifies particle propagation so trajectories better match the posterior and importance weights become less variable, but it is not standard controlled SMC: there is no Bellman equation, no explicit control cost plus terminal reward, no twisted path measure via optimal potentials, and no backward information recursion [1506.03338].

An earlier precursor is the SMC sampler with partial rejection control, which modifies the mutation step through a threshold \(c_t\) and induces a new effective forward kernel
\[
M_t^\ast(x_{t-1},x_t) = r(c_t,x_{t-1})^{-1} \min\left\{1,\, W_{t-1}(x_{t-1})\frac{w_t(x_{t-1},x_t)}{c_t} \right\} M_t(x_{t-1},x_t).
\]
The paper proves the stagewise variance inequality
\[
\operatorname{Var}_{M_t^\ast}\!\big[W_t^\ast(x_t)\big]\le \operatorname{Var}_{M_t}\!\big[W_t(x_t)\big],
\]
so it is a clear example of proposal-kernel modification aimed at reducing weight variance, although it does not formulate a global optimal control problem [0808.3466].

Another neighboring direction controls the resampling law rather than the proposal law. Independent resampling Sequential Monte Carlo replaces the usual dependent multinomial/bootstrap resampling by a construction that preserves the same one-particle marginal law as classical resampling, while removing the induced dependence between resampled particles [1607.05758]. In the static setting, the paper proves
\[
\mathbb E(\widehat\Theta_N^{\rm IS}) = \mathbb E(\widehat\Theta_{M_N}^{\rm SIR}) = \mathbb E(\widehat\Theta_{M_N}^{\rm I\mbox{-}SIR}),
\]
and
\[
\operatorname{var}(\widehat\Theta_{M_N}^{\rm SIR}) = \operatorname{var}(\widehat\Theta_{M_N}^{\rm I\mbox{-}SIR}) +\frac{M_N-1}{M_N}\operatorname{var}(\widehat\Theta_N^{\rm IS}),
\]
so removing coupling lowers variance while preserving the target-oriented marginal effect. This enlarges the broader design space of “controlled SMC” by showing that one can control the joint resampling law itself [1607.05758].

Rare-event SISR methods offer an even earlier control perspective. In large-deviation problems, the resampling weights are chosen so that
\[
w_t(\mathbf y_t)\propto \frac{q_t(y_t\mid \mathbf y_{t-1})}{\widetilde q_t(y_t\mid \mathbf y_{t-1})},
\]
and in the simplest exponential-tilting case
\[
w_t(Y_t)=e^{\theta_b \xi_t-\psi(\theta_b)}.
\]
The resulting estimators are shown to be logarithmically efficient, with criteria such as
\[
m\,\mathrm{Var}(\widehat\alpha_B)=p_n^2 e^{o(n)},
\]
so the control acts on particle allocation through resampling rather than on the proposal kernel itself [1202.4582].

Sequentially Constrained Monte Carlo (SCMC) is closer to tempered SMC samplers than to cSMC proper. It defines a sequence of targets by progressively enforcing a difficult constraint, for example
\[
\pi_t(\theta)\propto f(\theta)\|\theta\|_{\mathcal A}^{\tau_t},
\]
or, for monotone regression,
\[
\pi(\boldsymbol{\beta},\sigma^2\mid \mathbf X,\mathbf y,\tau) \propto \pi(\boldsymbol{\beta},\sigma^2)\, \mathcal N(\mathbf y-\mathbf X\boldsymbol\beta;\mathbf 0,\sigma^2\mathbf I)\prod_{i=1}^n \Phi\!\left(\tau\left.\frac{\partial \mathbf X\boldsymbol\beta}{\partial x}\right|_{x=x_i}\right).
\]
SCMC therefore contributes a path-design principle for constrained targets, but not a learned twisting or dynamic-programming control law [1410.8209].

## 5. Applications and empirical record

The original cSMC paper reports substantial gains over state-of-the-art methods at a fixed computational complexity on a variety of applications, including state-space models and complex static models [1708.08396]. Subsequent work specialized the framework to diffusion pseudolikelihoods under partial observation, bridge augmentation, and hypoelliptic dynamics [2507.14535], and to online filtering through rolling-window control adaptation [2508.00696].

| Setting | Baseline / configuration | Reported result |
|---|---|---|
| Neuroscience state-space model | PMMH ESS, BPF vs cSMC | BPF: \((4356,2442)\); cSMC: \((20973,13235)\) [1708.08396] |
| Log-Gaussian Cox process (\(d=900\)) | standard AIS, adaptive AIS | Variance of log-marginal-likelihood estimates \(573\times\) smaller than standard AIS, \(200\times\) smaller than adaptive AIS; MSE of adaptive AIS \(920\times\) larger than cSMC [1708.08396] |
| Partially observed hypoelliptic FHN | PMMH with cSMC (10 particles) vs BPF (125 particles) | cSMC was about twice as fast and had negligible likelihood-estimate dispersion compared with the BPF [2507.14535] |
| Online HMM inference | standard particle filtering approaches | Improved estimation accuracy and robustness in higher dimensions [2508.00696] |

The state-space results in the original cSMC paper are not limited to PMMH. In a low-noise, partially observed Lorenz-96 system, cSMC sharply reduces relative variance of log marginal-likelihood estimates across parameter settings, with strongest improvements in difficult regimes such as low observation noise, misspecified parameters, and higher dimensions [1708.08396]. In Bayesian logistic regression on three real datasets, cSMC achieves ESS near \(100\%\) and variance/RMSE reductions in \(\log Z\) estimation, often by several orders of magnitude [1708.08396]. In a classic nonlinear multimodal filtering benchmark, cSMC with RBF policy classes strongly improves ESS and reduces variance of \(\log Z\), especially at high signal-to-noise ratio [1708.08396].

The diffusion-inference literature shows that cSMC is particularly valuable when the likelihood itself is only available through a high-dimensional Feynman–Kac representation. In semi-linear SDEs discretized by splitting schemes, cSMC is used because naive bootstrap particle filtering can have enormous variance, especially when observations are informative or bridge dimensions are large [2507.14535]. Bridge augmentation reduces discretization bias and, under Assumption 2, the bridged transition converges in \(L_1\) to the true transition density as the number of bridge steps grows [2507.14535]. In practice, cSMC converged in about \(2\)–\(3\) iterations on bridged examples, around \(20\) particles per cSMC iteration were sufficient for stable estimates in SPSA measurements, and PMMH on the FHN example used \(10\) particles for cSMC versus \(125\) for BPF [2507.14535].

## 6. Limitations, misconceptions, and current directions

The formal analysis of cSMC is strong but not assumption-free. The theory relies on bounded potentials, suitable measurability and integrability, invertible Gram matrices for linear regression, smooth dependence of \(M_t^\psi(\exp(-\Phi^\top\beta))\) on \(\beta\), CLTs for particle empirical measures, and contraction or Lipschitz conditions for iterated ADP [1708.08396]. Practical failure modes include poor function classes, intractable twisted kernels, ill-conditioned regressions, approximation error accumulation in a single backward pass, and contraction assumptions that may fail globally [1708.08396].

Domain-specific applications introduce additional caveats. In diffusion inference, bridge augmentation still increases cost substantially; if the numerical scheme itself is unstable, cSMC can break because policy learning is corrupted; and Strang requires invertibility of \(\Gamma_\Delta\), sometimes only available after refining the time step via bridging [2507.14535]. ORCSMC is explicitly a finite-horizon online approximation to offline cSMC: it preserves bounded per-observation computation and memory by restricting adaptation to a rolling window, but therefore sacrifices global offline optimality [2508.00696].

A recurrent misconception is to identify cSMC with any adaptive SMC scheme. NASMC is proposal adaptation for SMC, not control-theoretic SMC; its alignment with cSMC is partial, with strong overlap in goal but not in formalism [1506.03338]. Independent resampling modifies the joint law of the resampling step while keeping the same single-particle marginal, so it is best viewed as control of offspring coupling rather than control of the proposal path law [1607.05758]. SCMC progressively imposes a difficult constraint by redesigning the intermediate targets, which makes it closely related to tempered SMC samplers rather than to optimal-control or twisting-based cSMC [1410.8209].

The contemporary literature therefore supports a layered interpretation. In the strict sense, controlled sequential Monte Carlo is the twisted Feynman–Kac, KL-control, Bellman-recursion framework introduced for general state-space and static models [1708.08396]. In a broader historical and methodological sense, the field also includes proposal adaptation, mutation-kernel control, resampling control, and sequential target design, all directed at the same central pathology: variance and degeneracy caused by mismatch between particle dynamics and the target distribution [1506.03338][0808.3466][1607.05758][1410.8209].

Source: https://www.emergentmind.com/topics/controlled-sequential-monte-carlo