Papers
Topics
Authors
Recent
Search
2000 character limit reached

Consistent inverse optimal control for infinite time-horizon discounted nonlinear systems under noisy observations

Published 8 Feb 2026 in math.OC | (2602.07874v1)

Abstract: Inverse optimal control (IOC) aims to estimate the underlying cost that governs the observed behavior of an expert system. However, in practical scenarios, the collected data is often corrupted by noise, which poses significant challenges for accurate cost function recovery. In this work, we propose an IOC framework that effectively addresses the presence of observation noise. In particular, compared to our previous work \cite{wang2025consistent}, we consider the case of discrete-time, infinite-horizon, discounted MDPs whose transition kernel is only weak Feller. By leveraging the occupation measure framework, we first establish the necessary and sufficient optimality conditions for the expert policy and then construct an infinite dimensional optimization problem based on these conditions. This problem is then approximated by polynomials to get a finite-dimensional numerically solvable one, which relies on the moments of the state-action trajectory's occupation measure. More specifically, the moments are robustly estimated from the noisy observations by a combined misspecified Generalized Method of Moments (GMM) estimator derived from observation model and system dynamics. Consequently, the entire algorithm is based on convex optimization which alleviates the issues that arise from local minima and is asymptotically and statistically consistent. Finally, the performance of the proposed method is illustrated through numerical examples.

Authors (3)

Summary

  • The paper develops a globally convex inverse optimal control program that recovers linear feature-based costs for infinite-horizon discounted systems with weak Feller dynamics, including deterministic systems.
  • It combines noise-induced and dynamics-induced moment conditions in a misspecified GMM estimator, with estimation bias controlled by polynomial approximation error rather than optimization local minima.
  • Experiments on linear and nonlinear systems show errors near 0.02 at 512 trajectories, while larger polynomial bases reduce approximation error once sufficient data are available.

Overview

This paper addresses inverse optimal control (IOC) for discrete-time, infinite-horizon, discounted Markov decision processes (MDPs) whose transition kernel is only weakly Feller, with demonstration data corrupted by additive observation noise. The goal is to recover the parameterization θ\theta_\ell of a linearly structured cost =θφ\ell = \theta_\ell^\top \varphi under which the expert's stationary Markov policy is optimal. The work extends the authors' prior contribution [(2602.07874) references wang2025consistent], which required a strongly Feller kernel; relaxing to weak Feller admits deterministic dynamics as a special case. The resulting algorithm is entirely convex — avoiding local minima — and is proven asymptotically and statistically consistent.

Problem formulation

The expert solves an infinite-horizon discounted optimal control problem over a compact Archimedean state space XX and state-independent action set AA, with transition kernel QQ satisfying only the weak Feller property (i.e., QQ pushes bounded measurable functions to bounded measurable functions). The cost belongs to a known continuous feature family ˉ=θφ\bar\ell = \theta_\ell^\top \varphi. Two regularity assumptions are imposed: the expert uses a stationary Markov policy, and there exist suboptimal state-action pairs of non-zero measure, which rules out the degenerate case θ=0\theta_\ell = 0 where every policy is optimal.

The key structural device is the occupation measure reformulation. The nonlinear forward problem is equivalently expressed as an infinite-dimensional linear program over occupation measures (μt)(\mu_t), via the flow constraint Pxμt+1=QμtP_x\mu_{t+1} = Q\mu_t. Strong duality holds between this primal LP and its dual over value functions =θφ\ell = \theta_\ell^\top \varphi0 satisfying the Bellman inequality =θφ\ell = \theta_\ell^\top \varphi1. Consequently, the KKT conditions are both necessary and sufficient for optimality, so the IOC problem reduces to finding a feasible =θφ\ell = \theta_\ell^\top \varphi2 pair whose complementary slackness holds against the expert's discounted occupation measure =θφ\ell = \theta_\ell^\top \varphi3.

A notable refinement over the authors' earlier work: Proposition 3 establishes that complementary slackness against the infinite-horizon measure is equivalent to complementary slackness against a finite-horizon normalized measure =θφ\ell = \theta_\ell^\top \varphi4 built from trajectories of length =θφ\ell = \theta_\ell^\top \varphi5, provided a persistent excitation condition holds (=θφ\ell = \theta_\ell^\top \varphi6 absolutely continuous w.r.t. =θφ\ell = \theta_\ell^\top \varphi7). Unlike the prior proposition requiring a deterministic expert policy, this argument relies only on the existence of a Radon–Nikodym derivative and shared conditional kernels, and therefore applies to any stochastic Markov policy.

Convex program and polynomial approximation

The IOC problem is cast as an infinite-dimensional convex program: minimize =θφ\ell = \theta_\ell^\top \varphi8 subject to =θφ\ell = \theta_\ell^\top \varphi9, an integral normalization constraint preventing trivial solutions, membership constraints on XX0 and XX1, and norm bounds. Proposition 4 shows this program attains minimum value zero, and its solutions characterize exactly those costs for which the expert policy is optimal. This equivalence means that solving the IOC problem requires only global optimization of a convex objective — local-minimum pathologies afflicting PMP-violation methods [rickenbach2024inverse, keshavarz2011imputing] do not arise here.

Finite-dimensional approximation restricts XX2 and XX3 to polynomial spaces of degrees XX4 and XX5, with the features XX6 and the pushed-forward basis XX7 approximated in a common polynomial basis via matrices XX8, XX9, AA0. The remaining unknown is the moment vector AA1, which must be estimated from noisy observations.

Moment estimation via misspecified GMM

Observations take the form AA2 with i.i.d., state-independent noise of known distribution AA3. Two families of moment conditions are derived:

  • Noise-induced moments: expanding AA4 by the multi-index binomial theorem yields a correctly specified moment condition relating observed discounted power sums to AA5 through an invertible lower-triangular matrix AA6.
  • Dynamics-induced moments: the occupation-measure evolution implies shifted-state moments satisfy AA7, a misspecified condition because it involves the polynomial approximation error of AA8.

These are combined into a two-step misspecified GMM estimator [hall2003large]. Because the model is misspecified, no weight matrix yields asymptotic efficiency; the paper instead adopts a block-diagonal weight matrix based on sample covariance estimates of the two moment blocks, plus regularization AA9. Beyond theoretical necessity, the block structure mitigates poor conditioning from high cross-correlation between the observation and dynamics moments. The closed-form weighted least-squares solution makes the estimator cheap and stable.

Consistency analysis

The consistency argument proceeds through three lemmas:

  1. Estimator convergence: QQ0 converges in probability to a limit QQ1 with distance from the true moment vector bounded by QQ2, i.e., the residual error is controlled purely by the polynomial approximation quality of the dynamics lifting.
  2. Error propagation: differences in optimal objective values between the true-moment and estimated-moment programs are bounded by QQ3, exploiting the norm constraints on the coefficients.
  3. Feasibility transfer: a constructive lemma shows any solution of the approximated program can be adjusted by a multiple of a known interior point QQ4 to remain feasible in the exact program, with objective degradation proportional to QQ5.

The main theorem combines these: under the triple limit of increasing trajectory count QQ6 and polynomial degrees (QQ7 innermost, then QQ8), plim of the objective at the estimated solution converges to zero — the optimal value of the exact infinite-dimensional program. Since zero objective value characterizes cost recovery (Proposition 4), this establishes asymptotic and statistical consistency. An implicit assumption worth noting: norm bounds QQ9 must be chosen large enough that they are inactive at optimum, yet small enough for the propagation bound to be informative; the proof assumes such bounds exist without quantifying them.

Numerical evaluation

Experiments cover a 2-D linear system solved via the algebraic Riccati equation and a nonlinear temperature control system (radiation/convection dynamics, fourth-order polynomial form) demonstrated by a discounted MPC controller. Non-negativity constraints are enforced via Weighted Sum-of-Squares relaxations, valid since QQ0 is Archimedean [putinar1993positive].

Key results:

Experiment Setting Result
Linear system QQ1, QQ2, QQ3, QQ4 Errors QQ5: mean QQ6, std QQ7
Temperature, noise/dataset study 400 trials, QQ8, QQ9 Mean error ˉ=θφ\bar\ell = \theta_\ell^\top \varphi0 at ˉ=θφ\bar\ell = \theta_\ell^\top \varphi1 across all noise levels
Temperature, degree study ˉ=θφ\bar\ell = \theta_\ell^\top \varphi2 Higher-degree bases reduce error once statistical noise no longer dominates

Two empirical observations align with theory: error decreases and plateaus with dataset size (the plateau reflects the approximation floor), and higher polynomial degrees lower that floor only when sufficient data is available for reliable moment estimation. At small datasets, all degree configurations perform comparably because statistical noise dominates — a practical caveat the paper acknowledges but does not quantify with sample-complexity bounds.

Limitations and open questions

Several assumptions constrain applicability. The noise must be additive, i.i.d., state-independent, and have a known distribution (from sensor calibration); correlated or multiplicative sensing errors fall outside the framework. The cost must be linear in known continuous features, and the state-action spaces must be compact and Archimedean. The persistent excitation assumption ties the required trajectory length ˉ=θφ\bar\ell = \theta_\ell^\top \varphi3 to the support of the infinite-horizon state marginal, which may be demanding for slowly mixing systems. The consistency guarantee is asymptotic in a specific nested limit order (ˉ=θφ\bar\ell = \theta_\ell^\top \varphi4, then ˉ=θφ\bar\ell = \theta_\ell^\top \varphi5, then ˉ=θφ\bar\ell = \theta_\ell^\top \varphi6) and provides no finite-sample rates; the constant ˉ=θφ\bar\ell = \theta_\ell^\top \varphi7 in Lemma 6 depends on unknown quantities (the true occupation measure and noise moments through ˉ=θφ\bar\ell = \theta_\ell^\top \varphi8), so it does not translate into an a priori accuracy certificate. Finally, whether the block-diagonal GMM weighting is near-optimal among tractable choices for the misspecified model remains unexamined, and the interaction between the SOS relaxation hierarchy and the moment estimation error is not analyzed.

Conclusion

The paper delivers a convex, provably consistent IOC method for infinite-horizon discounted nonlinear MDPs under weak Feller dynamics and noisy demonstrations. Its technical contributions are the extension of finite-horizon complementary-slackness equivalence to stochastic Markov policies, and a combined misspecified GMM estimator that fuses noise-model and dynamics-based moment conditions so that residual bias is governed solely by polynomial approximation error. The main theorem guarantees convergence of the recovered cost to the truth as data and approximation capacity grow, and experiments on linear and nonlinear plants corroborate the predicted behavior. Open issues center on finite-sample guarantees, relaxed noise models, and explicit selection of norm bounds and basis degrees.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.