- The paper develops a globally convex inverse optimal control program that recovers linear feature-based costs for infinite-horizon discounted systems with weak Feller dynamics, including deterministic systems.
- It combines noise-induced and dynamics-induced moment conditions in a misspecified GMM estimator, with estimation bias controlled by polynomial approximation error rather than optimization local minima.
- Experiments on linear and nonlinear systems show errors near 0.02 at 512 trajectories, while larger polynomial bases reduce approximation error once sufficient data are available.
Overview
This paper addresses inverse optimal control (IOC) for discrete-time, infinite-horizon, discounted Markov decision processes (MDPs) whose transition kernel is only weakly Feller, with demonstration data corrupted by additive observation noise. The goal is to recover the parameterization θℓ of a linearly structured cost ℓ=θℓ⊤φ under which the expert's stationary Markov policy is optimal. The work extends the authors' prior contribution [(2602.07874) references wang2025consistent], which required a strongly Feller kernel; relaxing to weak Feller admits deterministic dynamics as a special case. The resulting algorithm is entirely convex — avoiding local minima — and is proven asymptotically and statistically consistent.
The expert solves an infinite-horizon discounted optimal control problem over a compact Archimedean state space X and state-independent action set A, with transition kernel Q satisfying only the weak Feller property (i.e., Q pushes bounded measurable functions to bounded measurable functions). The cost belongs to a known continuous feature family ℓˉ=θℓ⊤φ. Two regularity assumptions are imposed: the expert uses a stationary Markov policy, and there exist suboptimal state-action pairs of non-zero measure, which rules out the degenerate case θℓ=0 where every policy is optimal.
The key structural device is the occupation measure reformulation. The nonlinear forward problem is equivalently expressed as an infinite-dimensional linear program over occupation measures (μt), via the flow constraint Pxμt+1=Qμt. Strong duality holds between this primal LP and its dual over value functions ℓ=θℓ⊤φ0 satisfying the Bellman inequality ℓ=θℓ⊤φ1. Consequently, the KKT conditions are both necessary and sufficient for optimality, so the IOC problem reduces to finding a feasible ℓ=θℓ⊤φ2 pair whose complementary slackness holds against the expert's discounted occupation measure ℓ=θℓ⊤φ3.
A notable refinement over the authors' earlier work: Proposition 3 establishes that complementary slackness against the infinite-horizon measure is equivalent to complementary slackness against a finite-horizon normalized measure ℓ=θℓ⊤φ4 built from trajectories of length ℓ=θℓ⊤φ5, provided a persistent excitation condition holds (ℓ=θℓ⊤φ6 absolutely continuous w.r.t. ℓ=θℓ⊤φ7). Unlike the prior proposition requiring a deterministic expert policy, this argument relies only on the existence of a Radon–Nikodym derivative and shared conditional kernels, and therefore applies to any stochastic Markov policy.
Convex program and polynomial approximation
The IOC problem is cast as an infinite-dimensional convex program: minimize ℓ=θℓ⊤φ8 subject to ℓ=θℓ⊤φ9, an integral normalization constraint preventing trivial solutions, membership constraints on X0 and X1, and norm bounds. Proposition 4 shows this program attains minimum value zero, and its solutions characterize exactly those costs for which the expert policy is optimal. This equivalence means that solving the IOC problem requires only global optimization of a convex objective — local-minimum pathologies afflicting PMP-violation methods [rickenbach2024inverse, keshavarz2011imputing] do not arise here.
Finite-dimensional approximation restricts X2 and X3 to polynomial spaces of degrees X4 and X5, with the features X6 and the pushed-forward basis X7 approximated in a common polynomial basis via matrices X8, X9, A0. The remaining unknown is the moment vector A1, which must be estimated from noisy observations.
Moment estimation via misspecified GMM
Observations take the form A2 with i.i.d., state-independent noise of known distribution A3. Two families of moment conditions are derived:
- Noise-induced moments: expanding A4 by the multi-index binomial theorem yields a correctly specified moment condition relating observed discounted power sums to A5 through an invertible lower-triangular matrix A6.
- Dynamics-induced moments: the occupation-measure evolution implies shifted-state moments satisfy A7, a misspecified condition because it involves the polynomial approximation error of A8.
These are combined into a two-step misspecified GMM estimator [hall2003large]. Because the model is misspecified, no weight matrix yields asymptotic efficiency; the paper instead adopts a block-diagonal weight matrix based on sample covariance estimates of the two moment blocks, plus regularization A9. Beyond theoretical necessity, the block structure mitigates poor conditioning from high cross-correlation between the observation and dynamics moments. The closed-form weighted least-squares solution makes the estimator cheap and stable.
Consistency analysis
The consistency argument proceeds through three lemmas:
- Estimator convergence: Q0 converges in probability to a limit Q1 with distance from the true moment vector bounded by Q2, i.e., the residual error is controlled purely by the polynomial approximation quality of the dynamics lifting.
- Error propagation: differences in optimal objective values between the true-moment and estimated-moment programs are bounded by Q3, exploiting the norm constraints on the coefficients.
- Feasibility transfer: a constructive lemma shows any solution of the approximated program can be adjusted by a multiple of a known interior point Q4 to remain feasible in the exact program, with objective degradation proportional to Q5.
The main theorem combines these: under the triple limit of increasing trajectory count Q6 and polynomial degrees (Q7 innermost, then Q8), plim of the objective at the estimated solution converges to zero — the optimal value of the exact infinite-dimensional program. Since zero objective value characterizes cost recovery (Proposition 4), this establishes asymptotic and statistical consistency. An implicit assumption worth noting: norm bounds Q9 must be chosen large enough that they are inactive at optimum, yet small enough for the propagation bound to be informative; the proof assumes such bounds exist without quantifying them.
Numerical evaluation
Experiments cover a 2-D linear system solved via the algebraic Riccati equation and a nonlinear temperature control system (radiation/convection dynamics, fourth-order polynomial form) demonstrated by a discounted MPC controller. Non-negativity constraints are enforced via Weighted Sum-of-Squares relaxations, valid since Q0 is Archimedean [putinar1993positive].
Key results:
| Experiment |
Setting |
Result |
| Linear system |
Q1, Q2, Q3, Q4 |
Errors Q5: mean Q6, std Q7 |
| Temperature, noise/dataset study |
400 trials, Q8, Q9 |
Mean error ℓˉ=θℓ⊤φ0 at ℓˉ=θℓ⊤φ1 across all noise levels |
| Temperature, degree study |
ℓˉ=θℓ⊤φ2 |
Higher-degree bases reduce error once statistical noise no longer dominates |
Two empirical observations align with theory: error decreases and plateaus with dataset size (the plateau reflects the approximation floor), and higher polynomial degrees lower that floor only when sufficient data is available for reliable moment estimation. At small datasets, all degree configurations perform comparably because statistical noise dominates — a practical caveat the paper acknowledges but does not quantify with sample-complexity bounds.
Limitations and open questions
Several assumptions constrain applicability. The noise must be additive, i.i.d., state-independent, and have a known distribution (from sensor calibration); correlated or multiplicative sensing errors fall outside the framework. The cost must be linear in known continuous features, and the state-action spaces must be compact and Archimedean. The persistent excitation assumption ties the required trajectory length ℓˉ=θℓ⊤φ3 to the support of the infinite-horizon state marginal, which may be demanding for slowly mixing systems. The consistency guarantee is asymptotic in a specific nested limit order (ℓˉ=θℓ⊤φ4, then ℓˉ=θℓ⊤φ5, then ℓˉ=θℓ⊤φ6) and provides no finite-sample rates; the constant ℓˉ=θℓ⊤φ7 in Lemma 6 depends on unknown quantities (the true occupation measure and noise moments through ℓˉ=θℓ⊤φ8), so it does not translate into an a priori accuracy certificate. Finally, whether the block-diagonal GMM weighting is near-optimal among tractable choices for the misspecified model remains unexamined, and the interaction between the SOS relaxation hierarchy and the moment estimation error is not analyzed.
Conclusion
The paper delivers a convex, provably consistent IOC method for infinite-horizon discounted nonlinear MDPs under weak Feller dynamics and noisy demonstrations. Its technical contributions are the extension of finite-horizon complementary-slackness equivalence to stochastic Markov policies, and a combined misspecified GMM estimator that fuses noise-model and dynamics-based moment conditions so that residual bias is governed solely by polynomial approximation error. The main theorem guarantees convergence of the recovered cost to the truth as data and approximation capacity grow, and experiments on linear and nonlinear plants corroborate the predicted behavior. Open issues center on finite-sample guarantees, relaxed noise models, and explicit selection of norm bounds and basis degrees.