- The paper establishes that nonlinear MPPI inherits a deterministic MPC policy’s contraction and Control Lyapunov Function certificate when its state-dependent approximation error satisfies an explicit small-gain condition.
- The analysis decomposes MPPI error into a temperature-dependent bias and an O(M^-1/2) Monte Carlo term, showing that lowering temperature—not sampling covariance alone—removes the intrinsic infinite-sample bias.
- The resulting finite-horizon, high-probability bound combines exponential decay with separate floors for MPPI error, Gaussian process noise, and confidence failure, while pendulum experiments confirm the guarantee but reveal substantial conservatism.
Overview and contribution
This paper, the second in a three-part series on closed-loop stability of Model Predictive Path Integral (MPPI) control (2607.06945), establishes a closed-loop stability certificate for MPPI applied to nonlinear stochastic discrete-time systems. The central claim is deliberately framed as a stability-inheritance theorem rather than an existence theorem: the authors do not show that MPPI stabilizes arbitrary nonlinear systems from first principles. Instead, they assume that a deterministic nonlinear MPC policy π∗(x)=[U∗(x)]0 exists whose disturbance-free closed loop is certified by a Control Lyapunov Function (CLF) terminal cost and a contraction metric, and prove that finite-sample MPPI inherits this stability certificate when its approximation error satisfies an explicit small-gain condition. The result should be read as the nonlinear counterpart of the companion LTI paper (Yoon et al., 4 Jul 2026), where the DARE/LQR structure provides the stabilizing reference in closed form.
The main guarantee is finite-horizon, high-probability localized mean practical stability. For any prescribed horizon T and confidence level δ, the trajectory remains in a compact sublevel set with probability at least 1−δ, and on that event the expected deviation from equilibrium decays exponentially up to three residual floors: an MPPI approximation floor, a Gaussian process-noise floor, and a bad-event confidence floor.
The system is xk+1=f(xk,uk)+wk with additive Gaussian process noise wk∼N(0,Σw), regulation to an equilibrium x∗=0, and a shared finite-horizon cost J(x,U) comprising a stage cost, a CLF terminal cost Vf(x)=x⊤Px, and nominal rollouts. MPPI samples perturbation sequences around a warm-started nominal sequence, weights them by exp(−J/λ), and applies the self-normalized weighted average of the first control block. Crucially, MPPI and the deterministic MPC share the same cost functional, so the comparison between the MPPI control and T0 is between a sampling-based implementation and the optimizer of the same objective — not between unrelated controllers.
Two structural certificates are imposed on the disturbance-free nominal closed loop: a globally valid CLF decrease condition (which eliminates the terminal constraint set required by the classical Mayne et al. framework), and contraction with rate T1 in a uniformly bounded Control Contraction Metric T2 satisfying T3. Because the noise is strictly additive, the variational dynamics are unaffected by T4, so no Itô-correction to the metric is needed; the authors note explicitly that multiplicative noise would require a Hessian correction.
Two-component approximation error
The analysis decomposes the per-step control error into an infinite-sample temperature bias plus a Monte Carlo term:
T5
Three lemmas support this bound. First, structural weight bounds (T6 a.s., since T7) and a uniform positive lower bound T8 on the expected self-normalization denominator hold on compact sets via dominated convergence. Second, Hoeffding and sub-Gaussian concentration give T9 with conditional probability at least δ0. Third, under a compact-set regularity assumption on the MPC optimizer (coercivity, global uniqueness, uniform positive-definite Hessian δ1 — not global convexity), a Laplace-principle argument shows the infinite-sample bias is bounded by δ2, where both constants vanish as δ3.
A notable negative result within Proposition 3: reducing the sampling covariance δ4 does not generally remove the infinite-sample bias, since the perturbation distribution collapses to zero and the update converges to the warm-start rather than the optimizer. Only lowering the temperature removes the bias floor. In the LTI/quadratic case the intrinsic residual δ5 vanishes identically because the Gibbs posterior is exactly Gaussian; for nonlinear costs it need not vanish even when the warm-start equals the optimizer.
Small-gain inheritance of contraction
The core robustness step is trajectory-level and requires no differentiation of the random MPPI policy. On the good event where the approximation bound holds, a triangle-inequality argument in the CCM geodesic distance yields
δ6
with small-gain function δ7. If δ8, then δ9: the state-dependent part of the MPPI error modifies the contraction rate, while the constant part becomes an additive practical-stability floor. This condition links sample size and temperature directly to closed-loop stability.
Finite-horizon localization and the main theorem
Because Gaussian process noise has unbounded support, no bounded sublevel set can be invariant almost surely over an infinite horizon — the authors state this plainly as the reason the certificate cannot be global. Lemma 5 instead proves that for any 1−δ0 and 1−δ1 there exists a radius 1−δ2 such that 1−δ3, justifying all compact-set constants over the prescribed horizon.
Combining these ingredients, the main theorem gives, for 1−δ4,
1−δ5
a class-1−δ6 decay term plus three interpretable floors: Monte Carlo/temperature bias, process noise, and bad-event contribution (controlled via Cauchy–Schwarz using a worst-case constant 1−δ7 built from actuator saturation 1−δ8). Corollaries restate the result conditionally on the localization event and in ISS form, and a noise-free corollary shows that in the ideal limit 1−δ9, xk+1=f(xk,uk)+wk0, xk+1=f(xk,uk)+wk1, xk+1=f(xk,uk)+wk2, exponential convergence inherited from the nominal MPC policy is recovered exactly.
An explicit design procedure follows: choose xk+1=f(xk,uk)+wk3 for xk+1=f(xk,uk)+wk4, choose xk+1=f(xk,uk)+wk5 to satisfy the small-gain condition, then take xk+1=f(xk,uk)+wk6. Two structural facts are worth highlighting: xk+1=f(xk,uk)+wk7 does not depend directly on xk+1=f(xk,uk)+wk8 (noise enters only through the irreducible stochastic floor and the localization radius), and improving the sampling covariance — e.g., via CoVO-MPC-style design — reduces xk+1=f(xk,uk)+wk9 by shrinking the concentration constant. The noise floor interfaces with the third paper in the series, where online covariance estimation tightens wk∼N(0,Σw)0 without changing wk∼N(0,Σw)1.
Numerical validation
Experiments use a discretized pendulum (wk∼N(0,Σw)2, damping wk∼N(0,Σw)3, wk∼N(0,Σw)4, horizon wk∼N(0,Σw)5, DARE terminal cost) retaining the full wk∼N(0,Σw)6 nonlinearity. The certified contraction rate is wk∼N(0,Σw)7 (empirical rate wk∼N(0,Σw)8), the estimated bias gain is wk∼N(0,Σw)9 with a constant temperature floor x∗=00, and the small-gain condition holds with margin x∗=01. Three findings stand out:
- One-step drift inequality: across 150 sampled states and x∗=02, the geodesic drift bound of the robustness proposition held at every tested state (0/150 violations, minimum slack 2.44).
- Localization: with x∗=03, x∗=04, the empirical quantile choice of x∗=05 achieved x∗=06 exactly, and x∗=07 remained stable across sample counts (x∗=08 for x∗=09).
- Three-floor bound: the certified mean bound held for all J(x,U)0 at every tested J(x,U)1, with the empirical decay rate tracking J(x,U)2 as J(x,U)3 grows.
The authors are candid that the certificate is valid but loose: the empirical localized mean converges to roughly J(x,U)4–J(x,U)5 while the certified floor is of order J(x,U)6, attributable to worst-case use of saturated control on the bad event and global Lipschitz propagation. They also concede that the small-gain condition is genuinely restrictive — configurations with weak damping or ill-conditioned metrics (J(x,U)7) yielded J(x,U)8 during calibration, in which case the theorem certifies nothing; the reported configuration was selected precisely for its large contraction margin.
Limitations and open questions
Several restrictions bound the scope of the result. The inheritance framing assumes existence of a stabilizing, contracting deterministic MPC policy with a unique nondegenerate optimizer on the operating set; verifying Assumption 5 (compact-set regularity) for a given nonlinear cost is itself nontrivial, and the authors note it is not automatic even with quadratic stage costs. Global Lipschitz dynamics and bounded applied control are assumed throughout. The guarantee is finite-horizon and high-probability rather than almost-surely invariant over infinite horizons — a necessary consequence of unbounded Gaussian noise, but a weaker statement than classical MPC stability. Finally, the certificate says nothing about tasks where the cost includes obstacle or other non-quadratic terms (the motivating coupled-pendulum experiment is explicitly excluded from any stability claim), and extending the small-gain condition via state-dependent metric bounds in place of global Lipschitz constants remains open.
Conclusion
The paper delivers a rigorous, honestly scoped stability certificate for nonlinear MPPI: an explicit small-gain condition under which a sampling-based implementation inherits the CLF-plus-contraction certificate of a stabilizing deterministic MPC policy, quantified through a two-component approximation error with an J(x,U)9 Monte Carlo term and a temperature-controlled bias floor. The resulting finite-horizon localized mean practical stability bound, its ISS restatement, and the concrete Vf(x)=x⊤Px0–Vf(x)=x⊤Px1–Vf(x)=x⊤Px2 design procedure connect sampling-based control theory to the classical Mayne et al. framework, while the numerical experiments confirm both the mechanism and the acknowledged conservatism of the sufficient conditions.