Papers
Topics
Authors
Recent
Search
2000 character limit reached

Stochastic Stability of Nonlinear MPPI via Contraction Theory and Control Lyapunov Functions

Published 8 Jul 2026 in eess.SY and math.OC | (2607.06945v1)

Abstract: Model Predictive Path Integral (MPPI) control is directly implementable on nonlinear systems because its online update requires only forward rollouts of the dynamics, not gradients, linearizations, or convex optimization. However, this algorithmic flexibility does not by itself provide a closed-loop stability certificate. This paper establishes such a certificate through a stability-inheritance argument. We assume that there exists a deterministic nonlinear MPC policy whose disturbance-free closed loop is certified by a Control Lyapunov Function terminal cost and a contraction metric, and we show that finite-sample MPPI inherits the nominal contraction when its sampling-based update approximates this reference policy with sufficient accuracy. The approximation error decomposes into a finite-temperature bias floor and a Monte Carlo term that vanishes at the inverse square-root rate in the sample count. Under an explicit small-gain condition, the resulting MPPI closed loop satisfies a finite-horizon, high-probability localized mean practical stability bound with residual floors due to MPPI approximation error, Gaussian process noise, and bad sampling events. The paper also gives an ISS-type restatement and a finite-horizon design procedure for choosing the localization set, temperature, and sample count.

Authors (2)

Summary

  • The paper establishes that nonlinear MPPI inherits a deterministic MPC policy’s contraction and Control Lyapunov Function certificate when its state-dependent approximation error satisfies an explicit small-gain condition.
  • The analysis decomposes MPPI error into a temperature-dependent bias and an O(M^-1/2) Monte Carlo term, showing that lowering temperature—not sampling covariance alone—removes the intrinsic infinite-sample bias.
  • The resulting finite-horizon, high-probability bound combines exponential decay with separate floors for MPPI error, Gaussian process noise, and confidence failure, while pendulum experiments confirm the guarantee but reveal substantial conservatism.

Overview and contribution

This paper, the second in a three-part series on closed-loop stability of Model Predictive Path Integral (MPPI) control (2607.06945), establishes a closed-loop stability certificate for MPPI applied to nonlinear stochastic discrete-time systems. The central claim is deliberately framed as a stability-inheritance theorem rather than an existence theorem: the authors do not show that MPPI stabilizes arbitrary nonlinear systems from first principles. Instead, they assume that a deterministic nonlinear MPC policy π(x)=[U(x)]0\pi^*(x)=[U^*(x)]_0 exists whose disturbance-free closed loop is certified by a Control Lyapunov Function (CLF) terminal cost and a contraction metric, and prove that finite-sample MPPI inherits this stability certificate when its approximation error satisfies an explicit small-gain condition. The result should be read as the nonlinear counterpart of the companion LTI paper (Yoon et al., 4 Jul 2026), where the DARE/LQR structure provides the stabilizing reference in closed form.

The main guarantee is finite-horizon, high-probability localized mean practical stability. For any prescribed horizon TT and confidence level δ\delta, the trajectory remains in a compact sublevel set with probability at least 1δ1-\delta, and on that event the expected deviation from equilibrium decays exponentially up to three residual floors: an MPPI approximation floor, a Gaussian process-noise floor, and a bad-event confidence floor.

Problem formulation

The system is xk+1=f(xk,uk)+wkx_{k+1}=f(x_k,u_k)+w_k with additive Gaussian process noise wkN(0,Σw)w_k\sim\mathcal{N}(0,\Sigma_w), regulation to an equilibrium x=0x^*=0, and a shared finite-horizon cost J(x,U)J(x,U) comprising a stage cost, a CLF terminal cost Vf(x)=xPxV_f(x)=x^\top P x, and nominal rollouts. MPPI samples perturbation sequences around a warm-started nominal sequence, weights them by exp(J/λ)\exp(-J/\lambda), and applies the self-normalized weighted average of the first control block. Crucially, MPPI and the deterministic MPC share the same cost functional, so the comparison between the MPPI control and TT0 is between a sampling-based implementation and the optimizer of the same objective — not between unrelated controllers.

Two structural certificates are imposed on the disturbance-free nominal closed loop: a globally valid CLF decrease condition (which eliminates the terminal constraint set required by the classical Mayne et al. framework), and contraction with rate TT1 in a uniformly bounded Control Contraction Metric TT2 satisfying TT3. Because the noise is strictly additive, the variational dynamics are unaffected by TT4, so no Itô-correction to the metric is needed; the authors note explicitly that multiplicative noise would require a Hessian correction.

Two-component approximation error

The analysis decomposes the per-step control error into an infinite-sample temperature bias plus a Monte Carlo term:

TT5

Three lemmas support this bound. First, structural weight bounds (TT6 a.s., since TT7) and a uniform positive lower bound TT8 on the expected self-normalization denominator hold on compact sets via dominated convergence. Second, Hoeffding and sub-Gaussian concentration give TT9 with conditional probability at least δ\delta0. Third, under a compact-set regularity assumption on the MPC optimizer (coercivity, global uniqueness, uniform positive-definite Hessian δ\delta1 — not global convexity), a Laplace-principle argument shows the infinite-sample bias is bounded by δ\delta2, where both constants vanish as δ\delta3.

A notable negative result within Proposition 3: reducing the sampling covariance δ\delta4 does not generally remove the infinite-sample bias, since the perturbation distribution collapses to zero and the update converges to the warm-start rather than the optimizer. Only lowering the temperature removes the bias floor. In the LTI/quadratic case the intrinsic residual δ\delta5 vanishes identically because the Gibbs posterior is exactly Gaussian; for nonlinear costs it need not vanish even when the warm-start equals the optimizer.

Small-gain inheritance of contraction

The core robustness step is trajectory-level and requires no differentiation of the random MPPI policy. On the good event where the approximation bound holds, a triangle-inequality argument in the CCM geodesic distance yields

δ\delta6

with small-gain function δ\delta7. If δ\delta8, then δ\delta9: the state-dependent part of the MPPI error modifies the contraction rate, while the constant part becomes an additive practical-stability floor. This condition links sample size and temperature directly to closed-loop stability.

Finite-horizon localization and the main theorem

Because Gaussian process noise has unbounded support, no bounded sublevel set can be invariant almost surely over an infinite horizon — the authors state this plainly as the reason the certificate cannot be global. Lemma 5 instead proves that for any 1δ1-\delta0 and 1δ1-\delta1 there exists a radius 1δ1-\delta2 such that 1δ1-\delta3, justifying all compact-set constants over the prescribed horizon.

Combining these ingredients, the main theorem gives, for 1δ1-\delta4,

1δ1-\delta5

a class-1δ1-\delta6 decay term plus three interpretable floors: Monte Carlo/temperature bias, process noise, and bad-event contribution (controlled via Cauchy–Schwarz using a worst-case constant 1δ1-\delta7 built from actuator saturation 1δ1-\delta8). Corollaries restate the result conditionally on the localization event and in ISS form, and a noise-free corollary shows that in the ideal limit 1δ1-\delta9, xk+1=f(xk,uk)+wkx_{k+1}=f(x_k,u_k)+w_k0, xk+1=f(xk,uk)+wkx_{k+1}=f(x_k,u_k)+w_k1, xk+1=f(xk,uk)+wkx_{k+1}=f(x_k,u_k)+w_k2, exponential convergence inherited from the nominal MPC policy is recovered exactly.

An explicit design procedure follows: choose xk+1=f(xk,uk)+wkx_{k+1}=f(x_k,u_k)+w_k3 for xk+1=f(xk,uk)+wkx_{k+1}=f(x_k,u_k)+w_k4, choose xk+1=f(xk,uk)+wkx_{k+1}=f(x_k,u_k)+w_k5 to satisfy the small-gain condition, then take xk+1=f(xk,uk)+wkx_{k+1}=f(x_k,u_k)+w_k6. Two structural facts are worth highlighting: xk+1=f(xk,uk)+wkx_{k+1}=f(x_k,u_k)+w_k7 does not depend directly on xk+1=f(xk,uk)+wkx_{k+1}=f(x_k,u_k)+w_k8 (noise enters only through the irreducible stochastic floor and the localization radius), and improving the sampling covariance — e.g., via CoVO-MPC-style design — reduces xk+1=f(xk,uk)+wkx_{k+1}=f(x_k,u_k)+w_k9 by shrinking the concentration constant. The noise floor interfaces with the third paper in the series, where online covariance estimation tightens wkN(0,Σw)w_k\sim\mathcal{N}(0,\Sigma_w)0 without changing wkN(0,Σw)w_k\sim\mathcal{N}(0,\Sigma_w)1.

Numerical validation

Experiments use a discretized pendulum (wkN(0,Σw)w_k\sim\mathcal{N}(0,\Sigma_w)2, damping wkN(0,Σw)w_k\sim\mathcal{N}(0,\Sigma_w)3, wkN(0,Σw)w_k\sim\mathcal{N}(0,\Sigma_w)4, horizon wkN(0,Σw)w_k\sim\mathcal{N}(0,\Sigma_w)5, DARE terminal cost) retaining the full wkN(0,Σw)w_k\sim\mathcal{N}(0,\Sigma_w)6 nonlinearity. The certified contraction rate is wkN(0,Σw)w_k\sim\mathcal{N}(0,\Sigma_w)7 (empirical rate wkN(0,Σw)w_k\sim\mathcal{N}(0,\Sigma_w)8), the estimated bias gain is wkN(0,Σw)w_k\sim\mathcal{N}(0,\Sigma_w)9 with a constant temperature floor x=0x^*=00, and the small-gain condition holds with margin x=0x^*=01. Three findings stand out:

  • One-step drift inequality: across 150 sampled states and x=0x^*=02, the geodesic drift bound of the robustness proposition held at every tested state (0/150 violations, minimum slack 2.44).
  • Localization: with x=0x^*=03, x=0x^*=04, the empirical quantile choice of x=0x^*=05 achieved x=0x^*=06 exactly, and x=0x^*=07 remained stable across sample counts (x=0x^*=08 for x=0x^*=09).
  • Three-floor bound: the certified mean bound held for all J(x,U)J(x,U)0 at every tested J(x,U)J(x,U)1, with the empirical decay rate tracking J(x,U)J(x,U)2 as J(x,U)J(x,U)3 grows.

The authors are candid that the certificate is valid but loose: the empirical localized mean converges to roughly J(x,U)J(x,U)4–J(x,U)J(x,U)5 while the certified floor is of order J(x,U)J(x,U)6, attributable to worst-case use of saturated control on the bad event and global Lipschitz propagation. They also concede that the small-gain condition is genuinely restrictive — configurations with weak damping or ill-conditioned metrics (J(x,U)J(x,U)7) yielded J(x,U)J(x,U)8 during calibration, in which case the theorem certifies nothing; the reported configuration was selected precisely for its large contraction margin.

Limitations and open questions

Several restrictions bound the scope of the result. The inheritance framing assumes existence of a stabilizing, contracting deterministic MPC policy with a unique nondegenerate optimizer on the operating set; verifying Assumption 5 (compact-set regularity) for a given nonlinear cost is itself nontrivial, and the authors note it is not automatic even with quadratic stage costs. Global Lipschitz dynamics and bounded applied control are assumed throughout. The guarantee is finite-horizon and high-probability rather than almost-surely invariant over infinite horizons — a necessary consequence of unbounded Gaussian noise, but a weaker statement than classical MPC stability. Finally, the certificate says nothing about tasks where the cost includes obstacle or other non-quadratic terms (the motivating coupled-pendulum experiment is explicitly excluded from any stability claim), and extending the small-gain condition via state-dependent metric bounds in place of global Lipschitz constants remains open.

Conclusion

The paper delivers a rigorous, honestly scoped stability certificate for nonlinear MPPI: an explicit small-gain condition under which a sampling-based implementation inherits the CLF-plus-contraction certificate of a stabilizing deterministic MPC policy, quantified through a two-component approximation error with an J(x,U)J(x,U)9 Monte Carlo term and a temperature-controlled bias floor. The resulting finite-horizon localized mean practical stability bound, its ISS restatement, and the concrete Vf(x)=xPxV_f(x)=x^\top P x0–Vf(x)=xPxV_f(x)=x^\top P x1–Vf(x)=xPxV_f(x)=x^\top P x2 design procedure connect sampling-based control theory to the classical Mayne et al. framework, while the numerical experiments confirm both the mechanism and the acknowledged conservatism of the sufficient conditions.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.