Papers
Topics
Authors
Recent
Search
2000 character limit reached

Learning to control switching nonlinear systems with Koopman operator regression

Published 13 Jul 2026 in math.OC, eess.SY, and stat.ML | (2607.11344v1)

Abstract: In this work, we consider the identification and control of nonlinear systems with finite action spaces. The unknown dynamics are estimated from finite samples with Koopman operator regression in a reproducing kernel Hilbert space, yielding a linear switching predictive model, the switches governed by the value of the control variable. In order to perform control in closed-loop, the learned dynamics are employed in an infinite-horizon optimal control problem with time-varying stage cost, which is solved by means of model predictive control. In a theoretical analysis, we derive learning rates for the Koopman dynamics approximation. We further quantify, under suitable assumptions, the sub-optimality of the model predictive control strategy, both in the case of exact Koopman dynamics, and in the case of learned ones. Numerical simulations on the Duffing oscillator complement our theoretical findings.

Summary

  • The paper introduces a Koopman operator regression framework that lifts nonlinear dynamics into an RKHS, enabling globally linear models for switching systems.
  • It employs Tikhonov-regularized nonparametric regression on snapshot data to accurately learn the linear operators associated with each control input.
  • Numerical experiments on the Duffing oscillator demonstrate that increasing the prediction horizon and training data enhances MPC performance and stabilization.

Koopman Operator Regression for Control of Switching Nonlinear Systems

System Identification via Koopman Operator Lifting

The paper "Learning to control switching nonlinear systems with Koopman operator regression" (2607.11344) formulates a principled approach to identification and control of nonlinear discrete-time systems with a finite control alphabet using Koopman operator theory. It exploits the linearization properties of the Koopman formalism, which lifts nonlinear state dynamics into a potentially infinite-dimensional space of observables, specifically an RKHS, permitting globally linear models even for complex nonlinear dynamics.

For a system xt+1=f(xt,ut)x_{t+1} = \mathfrak{f}(x_t, u_t) with ∣U∣<∞|\mathcal{U}|<\infty, the canonical RKHS embedding ψx\psi_x defines lifted dynamics governed by a family of linear Koopman operators Ku\mathcal{K}_u, yielding zt+1=Kutztz_{t+1} = \mathcal{K}_{u_t} z_t. Critically, for each control uu, the system switches between different linear operators, resulting in a linear switching system in the lifted space.

Koopman operator regression is employed to learn Ku\mathcal{K}_u from snapshot data. The procedure involves nonparametric regression in RKHS, with regularization to ensure numerical stability. Given i.i.d. samples xi∼ρux_i \sim \rho_u, the operator K^u\widehat{\mathcal{K}}_u is fit via Tikhonov-regularized minimization of empirical error, yielding a closed-form solution via the representer theorem. The construction avoids assumptions of ergodicity or stationarity of the underlying dynamics, permitting broader applicability.

Theoretical guarantees are established: under boundedness and source condition assumptions, the learned operators converge with sample complexity scaling as n−1/6n^{-1/6} in the optimal regime. These rates match known results in supervised learning with kernel methods, confirming the statistical efficiency of the approach.

Model Predictive Control with Koopman Models

The control objective is defined in the lifted space via an infinite-horizon stage cost, potentially time-varying to accommodate the peculiarities of finite control sets which may not permit classical asymptotic stability. With access to learned Koopman models, model predictive control (MPC) is applied, computing a receding horizon policy by minimizing the finite-horizon value function, exploiting the efficiency of linear operator composition in the RKHS.

Rigorous analysis is provided for both exact Koopman-MPC (where true operators are known) and approximate Koopman-MPC (where operators are learned from finite data). A sub-optimality gap is derived for both settings; for exact MPC, the error decays exponentially with horizon ∣U∣<∞|\mathcal{U}|<\infty0, confirming that short horizons suffice for near-optimal control under mild controllability assumptions. For approximate MPC, additional terms capture the effect of model mismatch, with the closed-loop cost upper bounded in terms of the statistical error bounds from learning theory, and quantifying robustness to operator estimation error.

Numerical Results: Duffing Oscillator Control

Empirical validation is performed on the Duffing oscillator, a canonical nonlinear benchmark. Koopman operators are learned using random Fourier features of the Gaussian kernel from snapshot pairs generated by the system. Multiple control sets are considered, including both symmetric (e.g., ∣U∣<∞|\mathcal{U}|<\infty1) and asymmetric (e.g., ∣U∣<∞|\mathcal{U}|<\infty2) actions.

Sample trajectories for constant controls reveal characteristic nonlinear behaviors: Figure 1

Figure 1

Figure 1

Figure 1

Figure 1: Direct trajectories of the Duffing oscillator for fixed ∣U∣<∞|\mathcal{U}|<\infty3, illustrating nonlinear response and initial conditions.

Learning performance is evaluated via a key performance indicator (KPI), measuring discounted distance from the origin over a simulation horizon. The results unequivocally show that increasing the predictive horizon ∣U∣<∞|\mathcal{U}|<\infty4 and increasing the number of training snapshots ∣U∣<∞|\mathcal{U}|<\infty5 improve closed-loop performance: Figure 2

Figure 2

Figure 2

Figure 2: Closed-loop trajectories for MPC with predictive horizon ∣U∣<∞|\mathcal{U}|<\infty6, revealing attractor behaviors and limited stabilization.

For asymmetric control sets, the inclusion of costly high-magnitude controls accelerates approach to sliding surfaces and improves stabilization speed, at the expense of control cost: Figure 3

Figure 3

Figure 3

Figure 3: Duffing oscillator trajectories for fixed ∣U∣<∞|\mathcal{U}|<\infty7, demonstrating rapid state evolution under strong input.

Mean and standard deviation over multiple randomized initializations corroborate the theoretical claims. Notably, for ∣U∣<∞|\mathcal{U}|<\infty8 and ∣U∣<∞|\mathcal{U}|<\infty9, the system consistently stabilizes at the origin or within its neighborhood, confirming the effectiveness of the Koopman-MPC pipeline.

Implications and Future Directions

The methodology provides a coherent pipeline for data-driven control of nonlinear systems with finite action sets, circumventing the need for explicit nonlinear system identification. The decoupling of system identification and control via operator regression allows for modular analysis of error propagation, and the RKHS formalism enables spectral learning rates without geometric or ergodicity constraints.

The theoretical performance bounds connect statistical learning theory to robust closed-loop control guarantees, a crucial advance for principled data-driven reinforcement learning, especially in settings where model uncertainty is prominent. Practical implications include applicability to power electronics and hybrid systems with discrete controls, and for AI-driven control architectures requiring interpretable and tractable model representations.

Potential future developments include extending the learning theory to non-iid data, incorporating terminal ingredients in MPC for recursive feasibility, and enabling quasi-real time optimization for large-scale combinatorial action spaces. Application to real systems, integration with density estimation for sampling, and adaptation to continuous control alphabets via operator interpolation are promising research avenues.

Conclusion

This paper establishes a rigorous foundation for learning to control switching nonlinear systems using Koopman operator regression in an RKHS. The approach unifies system identification and control in the lifted space, delivers statistical learning guarantees and robust closed-loop analysis, and demonstrates empirical efficacy on challenging nonlinear benchmarks. It opens the path to scalable and theoretically principled data-driven control frameworks for hybrid and nonlinear systems.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.