- The paper provides the first explicit finite-sample closed-loop stability certificate for MPPI in LTI systems with quadratic cost.
- It decomposes the MPPI approximation error into a Monte Carlo term and a temperature-induced bias, quantifying each component's impact.
- The study derives computable sample-complexity thresholds and uses a Lyapunov-based argument to ensure exponential decay of the expected state norm.
Finite-Sample Stability Guarantees for MPPI in LTI Systems
Introduction and Context
Model Predictive Path Integral (MPPI) control is a sampling-based, gradient-free receding-horizon control method, widely adopted in robotics and autonomous systems due to its parallelizability and suitability for non-smooth, nonlinear dynamics. However, theoretical guarantees on the closed-loop stability of MPPI under finite-sample regimes, particularly in the presence of persistent process noise, have been largely absent in prior literature. This paper provides the first explicit finite-sample closed-loop stability certificate for MPPI applied to discrete-time linear time-invariant (LTI) systems with quadratic cost and additive Gaussian disturbances (2607.04006).
The target system is a discrete-time LTI plant
xk+1​=Axk​+Buk​+wk​,
where xk​∈Rn, uk​∈Rm, and wk​ is i.i.d. Gaussian noise. The cost is standard finite-horizon LQR with DARE terminal cost:
J(xk​,U)=i=0∑N−1​(xk+i⊤​Qxk+i​+uk+i⊤​Ruk+i​)+xk+N⊤​Pxk+N​
with P as the stabilizing DARE solution. This structure ensures that, for unconstrained problems, the first action of the finite-horizon optimizer coincides with the infinite-horizon LQR solution, for any planning horizon N≥1.
MPPI computes the control by
uk​=uˉ0​+∑j=1M​w(j)∑j=1M​w(j)ϵ0(j)​​
where M is the number of sampled noisy control sequences, and w(j) are importance weights from a Gibbs distribution parameterized by temperature xk​∈Rn0.
Main Results
Decomposition of MPPI Approximation Error
A central contribution is the explicit finite-sample bound for the deviation of the MPPI control from the LQR-optimal action. The error admits a two-term decomposition:
- Monte Carlo Term: xk​∈Rn1, representing the stochastic error from finite sampling, vanishing as xk​∈Rn2 increases.
- Temperature Bias: A nonzero bias emerges at finite xk​∈Rn3, vanishing only as xk​∈Rn4.
The closed-form characterization of the bias is given in terms of the cost matrices and MPPI temperature, with xk​∈Rn5 quantifying the mismatch between the infinite-sample MPPI mean and the LQR solution.
Lyapunov-Based Stability Certificate
The authors employ a Lyapunov perturbation argument, showing that when the MPPI error is sufficiently small, the stochastic Lyapunov decrease induced by the nominal LQR law absorbs the MPPI perturbation. This leads to a practical exponential stability result: on sample paths that remain inside a prescribed Lyapunov sublevel set xk​∈Rn6 up to time xk​∈Rn7, the expected state norm decays exponentially up to an explicit residual floor. The residual is composed of three terms:
- Process Noise Floor: Directly stemming from persistent Gaussian disturbances.
- MPPI Approximation Floor: Due to both Monte Carlo error and persistent temperature bias.
- Confidence Floor: Resulting from per-step sampling failure probability.
Explicit Finite-Sample and Sample-Complexity Bounds
The main theorems yield explicit, computable sample size thresholds xk​∈Rn8, depending on the DARE solution, system dimensions, temperature xk​∈Rn9, and user-supplied confidence/accuracy requirements. Certifiable stability (in expectation over process noise and sampling events) is achieved for all uk​∈Rm0.
The exponential stability result is further recast into a practical ISS bound, explicitly exhibiting the influence of process disturbance, MPPI error, and sampling confidence on the steady-state state norm in expectation. This frames the result within the established robust MPC literature.
Numerical Results and Empirical Validation
Simulations on a double-integrator benchmark validate the sufficiency and conservatism of the uk​∈Rm1 sample threshold. Empirical system stability emerges for sample counts significantly below uk​∈Rm2, demonstrating that the bound is sufficient but not tight. The stability decay rate bound is conservative compared to the actual LQR closed-loop rate, reflecting worst-case norm inequalities in the theoretical analysis.
Other key observations:
- The sample threshold is largely independent of the process noise variance, affecting only the noise floor, not the control-oriented stability margin.
- The effective sample size (ESS) is not a reliable real-time stability indicator for MPPI, as stable behavior persists even with low normalized ESS.
Theoretical and Practical Implications
This work closes an open gap in the theory of sampling-based stochastic control: explicit conditions under which receding-horizon MPPI yields exponentially stable closed-loop behavior for LTI/quadratic systems, under finite sampling. By tightly linking the finite-sample MPPI error structure to classical LQR Lyapunov margins, the analysis establishes a framework for quantifying and trading off compute (sample budget) versus control performance in practical deployments.
Strong claims:
- The stability guarantee is explicit, with all constants computable from system data and MPPI parameters.
- In the joint limit as uk​∈Rm3 and uk​∈Rm4, the results recover the standard stochastic LQR stability certificate.
- The bias due to finite sampling covariance cannot be eliminated by simply reducing uk​∈Rm5; only lowering uk​∈Rm6 (temperature) attenuates the bias.
Contradictory to prior assumptions:
- Prior works focused on optimizer convergence or single-step performance, whereas this analysis demonstrates the necessity of considering long-term closed-loop properties and their dependence on sampling and temperature.
Directions for Future Development
The restriction to unconstrained LTI/quadratic problems leaves several important extensions open:
- Constraints and Recursive Feasibility: Adapting the approach to constrained MPC would demand the integration of terminal set, recursive feasibility, and chance-constrained analysis, accounting for the inherent possibility of constraint violation under additive stochastic noise.
- Horizon-Uniformity: While the nominal LQR action is horizon-invariant, the MPPI sample complexity constants are not. Uniform sample-complexity guarantees as uk​∈Rm7 would require further technical advances.
- Nonlinear Systems: The conceptual framework is extendable to nonlinear dynamics via contraction metrics and CLF methods, but additional structural properties must be exploited.
Conclusion
This paper establishes, for the first time, explicit finite-sample closed-loop stability guarantees for MPPI applied to discrete-time LTI systems. The analysis quantifies the Monte Carlo and temperature-induced bias, connects sampling-based control to Lyapunov-based MPC stability theory, and produces readily computable sample-complexity thresholds. The theoretical results are empirically supported, and the decomposition of the residual floors provides actionable insight into regime selection for robust, stable control under sampling constraints. This work lays the groundwork for practical, stable applications of MPPI in stochastic, computationally constrained control environments, and points toward a robust framework for analysis beyond the linear-quadratic setting.