Deterministic Sampling CEM (dsCEM)
- The paper introduces dsCEM, which deterministically samples control trajectories using localized cumulative distributions to enhance performance in nonlinear MPC tasks.
- It leverages affine transformations of precomputed Dirac mixtures to replace Gaussian random sampling, ensuring smoother control inputs and efficient computation.
- Experimental results on benchmark tasks like Mountain Car and Cart-Pole show up to 30% cost reduction and 40% smoother controls in low-sample regimes.
Deterministic Sampling CEM (dsCEM) is a model predictive control (MPC) framework for nonlinear optimal control tasks that replaces the conventional random sampling step of the Cross-Entropy Method (CEM) with a deterministic procedure. This deterministic approach uses sample sets derived from localized cumulative distributions (LCDs), providing improved sample efficiency and smoother control input trajectories, particularly in regimes with a low number of samples. The dsCEM methodology preserves the core structure and tuning of standard CEM-MPC, enabling straightforward integration as a drop-in replacement for random sampling-based controllers (Walker et al., 7 Oct 2025).
1. Background: Standard CEM–MPC Formulation
In CEM–MPC, a finite-horizon control sequence $\vu_{k:k+N-1}$ is optimized to minimize a cost function for a nonlinear discrete-time system
$\vx_{n+1} = \va_n(\vx_n, \vu_n)\,,$
with an associated cumulative cost
$J_k(\vu_{k:k+N-1}) = g_N(\vx_{k+N}) + \sum_{n=k}^{k+N-1} g_n(\vx_n, \vu_n)\,.$
CEM places a Gaussian proposal distribution
$f(\vu;\vtheta_j) = \mathcal{N}(\vu ; \vmu_j, \mC_j), \quad \vtheta_j = (\vmu_j, \mC_j)$
over the flattened trajectory vector $\vu = [\vu_k^\top, \ldots, \vu_{k+N-1}^\top]^\top$. At each iteration, samples are drawn, their costs evaluated, the elite set (top samples) is selected, and the mean and covariance are updated: $\vmu_{j+1} = \frac{1}{|\mathcal{E}_j|} \sum_{\vu \in \mathcal{E}_j} \vu\,, \quad \mC_{j+1} = \frac{1}{|\mathcal{E}_j|} \sum_{\vu \in \mathcal{E}_j} (\vu - \vmu_{j+1})(\vu - \vmu_{j+1})^\top\,.$ This process iterates, returning either the first control from $\vmu_J$ or the best sampled sequence.
2. Deterministic Sampling via Localized Cumulative Distributions (LCDs)
The core innovation of dsCEM is the replacement of random Gaussian samples by a deterministic Dirac mixture, optimized according to LCD-based discrepancy criteria. For a $\vx_{n+1} = \va_n(\vx_n, \vu_n)\,,$0-dimensional PDF $\vx_{n+1} = \va_n(\vx_n, \vu_n)\,,$1, one defines the LCD at location $\vx_{n+1} = \va_n(\vx_n, \vu_n)\,,$2 and scale $\vx_{n+1} = \va_n(\vx_n, \vu_n)\,,$3 as
$\vx_{n+1} = \va_n(\vx_n, \vu_n)\,,$4
The Cramér–von Mises–type distance,
$\vx_{n+1} = \va_n(\vx_n, \vu_n)\,,$5
with $\vx_{n+1} = \va_n(\vx_n, \vu_n)\,,$6, quantifies the discrepancy of a Dirac mixture $\vx_{n+1} = \va_n(\vx_n, \vu_n)\,,$7 from $\vx_{n+1} = \va_n(\vx_n, \vu_n)\,,$8. Minimizing this discrepancy yields an optimal low-discrepancy sample skeleton $\vx_{n+1} = \va_n(\vx_n, \vu_n)\,,$9 for the isotropic Gaussian $J_k(\vu_{k:k+N-1}) = g_N(\vx_{k+N}) + \sum_{n=k}^{k+N-1} g_n(\vx_n, \vu_n)\,.$0, computed offline. For each CEM iteration (online), these samples are affinely transformed according to the current proposal parameters: $J_k(\vu_{k:k+N-1}) = g_N(\vx_{k+N}) + \sum_{n=k}^{k+N-1} g_n(\vx_n, \vu_n)\,.$1 ensuring the sample set deterministically matches the intended Gaussian.
3. Modular Variants: Sample Variability and Temporal Correlation
To avoid repeatability and enhance exploration while preserving determinism, dsCEM introduces three sample variability schemes and two temporal-correlation methods:
- Sample-set variability:
- V1 (Random Rotation): Applies a random rotation $J_k(\vu_{k:k+N-1}) = g_N(\vx_{k+N}) + \sum_{n=k}^{k+N-1} g_n(\vx_n, \vu_n)\,.$2 before affine transformation.
- V2 (Deterministic Joint-Density): Precomputes a block in $J_k(\vu_{k:k+N-1}) = g_N(\vx_{k+N}) + \sum_{n=k}^{k+N-1} g_n(\vx_n, \vu_n)\,.$3 and slices a fresh block per CEM iteration.
- V3 (Time-step Rotation): Applies a fixed random rotation per MPC step, subsequently using V2’s slicing.
- Temporal correlations:
- M1 (Fixed Correlation + Adaptive Variance): $J_k(\vu_{k:k+N-1}) = g_N(\vx_{k+N}) + \sum_{n=k}^{k+N-1} g_n(\vx_n, \vu_n)\,.$4. Here, $J_k(\vu_{k:k+N-1}) = g_N(\vx_{k+N}) + \sum_{n=k}^{k+N-1} g_n(\vx_n, \vu_n)\,.$5 is a fixed Toeplitz matrix structured according to a chosen $J_k(\vu_{k:k+N-1}) = g_N(\vx_{k+N}) + \sum_{n=k}^{k+N-1} g_n(\vx_n, \vu_n)\,.$6 power spectral density (PSD), with variances $J_k(\vu_{k:k+N-1}) = g_N(\vx_{k+N}) + \sum_{n=k}^{k+N-1} g_n(\vx_n, \vu_n)\,.$7 updated adaptively based on elite sets.
- M2 (Adaptive Full Covariance): Initializes from a colored-noise PSD and updates the full covariance following the elite-sample statistics.
The overall dsCEM-MPC algorithm incorporates these choices modularly, allowing flexible adaptation to problem structure and smoothness requirements.
4. Integration into Existing CEM Controllers
dsCEM is designed for compatibility with existing CEM-MPC implementations. Integration consists of substituting the standard random sampling with the deterministic LCD-based sample generation:
- Replace random sampling from $J_k(\vu_{k:k+N-1}) = g_N(\vx_{k+N}) + \sum_{n=k}^{k+N-1} g_n(\vx_n, \vu_n)\,.$8 with transforming a precomputed skeleton:
- Optional sample rotation via $J_k(\vu_{k:k+N-1}) = g_N(\vx_{k+N}) + \sum_{n=k}^{k+N-1} g_n(\vx_n, \vu_n)\,.$9
- Affine transformation as above
Additional controller parameters include:
- Selection of variability scheme (V1, V2, V3)
- Selection of temporal-correlation method (M1, M2)
- Precomputed pool of skeleton samples of desired size $f(\vu;\vtheta_j) = \mathcal{N}(\vu ; \vmu_j, \mC_j), \quad \vtheta_j = (\vmu_j, \mC_j)$0
The remainder of the CEM update, elite selection, iteration cycling, and hyperparameter tuning are unaltered.
5. Experimental Evaluation and Performance Metrics
dsCEM was evaluated on two canonical nonlinear MPC benchmarks:
- Mountain Car: RK4 discretization ($f(\vu;\vtheta_j) = \mathcal{N}(\vu ; \vmu_j, \mC_j), \quad \vtheta_j = (\vmu_j, \mC_j)$1, horizon $f(\vu;\vtheta_j) = \mathcal{N}(\vu ; \vmu_j, \mC_j), \quad \vtheta_j = (\vmu_j, \mC_j)$2), quadratic state/control cost, mild process noise.
- Cart-Pole Swing-Up: RK4 discretization ($f(\vu;\vtheta_j) = \mathcal{N}(\vu ; \vmu_j, \mC_j), \quad \vtheta_j = (\vmu_j, \mC_j)$3, horizon $f(\vu;\vtheta_j) = \mathcal{N}(\vu ; \vmu_j, \mC_j), \quad \vtheta_j = (\vmu_j, \mC_j)$4), standard nonlinear cart-pole states, augmented with trigonometric state components, and low process noise.
Both experiments used 3 CEM iterations, momentum $f(\vu;\vtheta_j) = \mathcal{N}(\vu ; \vmu_j, \mC_j), \quad \vtheta_j = (\vmu_j, \mC_j)$5, and 100 independent trials across sample sizes $f(\vu;\vtheta_j) = \mathcal{N}(\vu ; \vmu_j, \mC_j), \quad \vtheta_j = (\vmu_j, \mC_j)$6. Two performance metrics were assessed:
- Cumulative cost: $f(\vu;\vtheta_j) = \mathcal{N}(\vu ; \vmu_j, \mC_j), \quad \vtheta_j = (\vmu_j, \mC_j)$7
- Smoothness of control input: $f(\vu;\vtheta_j) = \mathcal{N}(\vu ; \vmu_j, \mC_j), \quad \vtheta_j = (\vmu_j, \mC_j)$8
6. Results, Guidelines, and Limitations
Quantitative Findings:
In the low-sample regime ($f(\vu;\vtheta_j) = \mathcal{N}(\vu ; \vmu_j, \mC_j), \quad \vtheta_j = (\vmu_j, \mC_j)$9), dsCEM-V2 achieved up to 30 % lower cumulative cost and up to 40 % lower smoothness score $\vu = [\vu_k^\top, \ldots, \vu_{k+N-1}^\top]^\top$0 compared to iCEM. All dsCEM variants required fewer iterations and samples to reach a given cost level. For example, with $\vu = [\vu_k^\top, \ldots, \vu_{k+N-1}^\top]^\top$1, dsCEM-V2 matched or exceeded the smoothness of iCEM run with $\vu = [\vu_k^\top, \ldots, \vu_{k+N-1}^\top]^\top$2 samples. The M2 method (full covariance) was less robust at very small $\vu = [\vu_k^\top, \ldots, \vu_{k+N-1}^\top]^\top$3, due to its elevated elite-set requirements. No formal hypothesis testing was reported, but interquartile ranges over 100 runs indicated non-overlapping medians in low-sample settings (Walker et al., 7 Oct 2025).
Hyperparameter Guidelines:
- Sample size $\vu = [\vu_k^\top, \ldots, \vu_{k+N-1}^\top]^\top$4: dsCEM provides greatest benefit for $\vu = [\vu_k^\top, \ldots, \vu_{k+N-1}^\top]^\top$5; above $\vu = [\vu_k^\top, \ldots, \vu_{k+N-1}^\top]^\top$6, differences vs random sampling diminish.
- Variability scheme: V2 is recommended for deterministic, high-smoothness control; V1 or V3 for some randomness.
- LCD pool size: Set equal to the largest $\vu = [\vu_k^\top, \ldots, \vu_{k+N-1}^\top]^\top$7 anticipated during control; computed offline using established LCD solvers.
- PSD exponent $\vu = [\vu_k^\top, \ldots, \vu_{k+N-1}^\top]^\top$8: Tune for smoothness; e.g., $\vu = [\vu_k^\top, \ldots, \vu_{k+N-1}^\top]^\top$9 yields "pink noise."
- Warm-start and momentum 0: 1.
Limitations and Extensions:
- M2 covariance adaptation is less effective in sample-scarce settings due to large 2.
- LCD-based affine transforms do not preserve LCD-optimality for non-isotropic proposals.
- Offline sample computation is computationally intensive for large 3.
- Potential future work includes integration with learned warm-starts (e.g., normalizing flows), adaptive online LCD pools for non-Gaussian proposals, theoretical analysis of affine-transformed sample discrepancy, and explicit extension to stochastic or chance-constrained MPC.
7. Significance and Prospective Directions
dsCEM introduces determinism and improved structure to the CEM-MPC sampling process, leading to enhanced sample efficiency and control smoothness, features especially pronounced in resource-limited regimes. The approach maintains compatibility with standard CEM controller architectures, requiring minimal interface adjustment. Extensions suggested include hybridization with learning-based warm-starts, further adaptation to structured noise models, and rigorous convergence theory. These directions may strengthen practical applicability in high-dimensional control, safety-critical planning, and non-Gaussian or uncertainty-aware MPC contexts (Walker et al., 7 Oct 2025).