Papers
Topics
Authors
Recent
Search
2000 character limit reached

Data-driven Linear Quadratic Integral Control: A Convex Formulation and Policy Gradient Approach

Published 16 Apr 2026 in eess.SY | (2604.14905v1)

Abstract: This paper studies the data-driven synthesis of linear quadratic integral (LQI) controllers for continuous-time systems. The objective is to achieve optimal state-feedback control with integral action for reference tracking using only measured data. To this end, we derive a data-driven closed-loop parameterization of the augmented dynamics that incorporates the integral state while relying solely on input-state-output measurements of the underlying system. Based on this parameterization, a data-driven convex optimization problem is formulated whose solution yields the optimal linear quadratic regulator (LQR) feedback gain for the augmented system without explicit knowledge of the system matrices. In addition, a policy gradient flow is derived to compute the optimal controller within the space of stabilizing gains. The proposed approach enables data-driven optimal tracking control while avoiding explicit state augmentation in the data collection phase. The effectiveness of the method is demonstrated through a numerical example involving a distributed generation unit (DGU) in a DC microgrid.

Summary

  • The paper introduces a data-driven LQI controller design that synthesizes optimal gains directly from empirical I/O data, eliminating the need for explicit model identification.
  • It formulates the controller synthesis as a convex SDP and employs a policy gradient flow that guarantees global convergence within the feasible set.
  • Empirical validation on a DC microgrid demonstrates robust voltage regulation with tracking errors as low as 4.3×10⁻⁴ in the Frobenius norm.

Data-driven LQI Control: Convex Synthesis and Policy Gradient Methods

Introduction and Problem Context

The paper "Data-driven Linear Quadratic Integral Control: A Convex Formulation and Policy Gradient Approach" (2604.14905) contributes to the current landscape of data-driven control by systematically addressing optimal state-feedback tracking via LQI control for continuous-time systems using measured I/O/state samples only. Conventional LQI design hinges on access to system matrices for augmented dynamics; the novelty here is a formulation that enables LQI controller synthesis directly from system data, sidestepping explicit model identification and state augmentation during data collection.

Theoretical Foundations

The approach is grounded in established linear systems and LQR theory, where the LQI framework extends standard state regulation to reference tracking via integral action. The paper first delineates the control landscape: full-state LQR, its convex and policy-gradient-based formulations, and the augmentation to LQI for tracking. The authors formalize the closed-loop dynamics of the augmented system, derive the stabilizability and detectability conditions, and clarify the structural distinctions between LQI and fixed-structure PI(D) controllers.

A pivotal observation is that while PID controller gain extraction via LQR on the augmented system is, in general, nonconvex and may not always be feasible due to structural constraints, LQI enjoys full-state feedback flexibility and conducive convexity properties.

Data-driven Closed-loop Parameterization

A core result is the derivation of a closed-loop parameterization for the augmented system utilizing only collected data. Following the sample covariance-based data representations from recent literature, the paper extends these methods to cater for the integral state within the augmented model, yet shows that the integral state itself need not be measured—the original system's I/O/state data suffices. The result is a full characterization of the augmented closed-loop via data matrices, facilitating subsequent control synthesis.

The theoretical guarantees rest on the rank conditions of the empirical data matrices (specifically, the concatenation [U X]\left[\begin{smallmatrix} \overline{U} \ \overline{X} \end{smallmatrix}\right]), which are satisfied under persistently exciting input sequences. This point ensures the practical implementability of the results.

Data-driven Convex Synthesis and Policy Gradient Flow

Given the data-driven representation, the paper develops a convex SDP formulation for LQI controller synthesis, with the controller gain directly parametrized by the solution to an optimization problem based on the collected trajectories (not requiring explicit system identification). The control optimality and closed-loop stability are guaranteed by constraints expressed in terms of data covariances and Schur complements.

In parallel, the authors formulate a policy gradient flow, projected onto the feasible set defined by linear equation constraints from the empirical data. This is achieved by deriving the analytical gradient of the LQR objective with respect to the data-driven parameters and projecting it onto the tangent space using the pseudoinverse of the sample covariance matrix. The gradient flow exhibits global convergence properties within the feasible set, thereby offering an online, iterative alternative to one-shot convex optimization.

Empirical Validation

The empirical section illustrates the efficacy of the proposed approaches on a Distributed Generation Unit (DGU) within a DC microgrid. The system is governed by an RLC filter and a time-varying static load. Data is generated in open-loop using persistently exciting inputs, and state-output samples are used for controller identification. Key experimental steps include:

  • Open-loop data collection with T=10T=10 samples.
  • Construction of the sample covariance matrices for the policy parameterization.
  • Synthesis of the optimal gain using the proposed SDP formulation.
  • Execution of the projected policy gradient initialized with diverse stabilizing gains.

The closed-loop application of the computed gain demonstrates robust regulation and sharp reference tracking, closely matching the gain computed via conventional model-based LQI design (with absolute error KKF=4.3×104\Vert K^* - K^\star \Vert_\mathrm{F} = 4.3 \times 10^{-4}).

Figure 1

Figure 2: Trajectories of the bus voltage v(t)v(t) and input u(t)u(t) during open-loop data collection and closed-loop control.

Convergence of the policy gradient flow is numerically validated for multiple initializations, with rapid residual decay to zero, corroborating the theoretical prediction of global convergence for all stabilizing initial feedbacks.

Figure 3

Figure 1: Normalized residuals K(t)KFK(0)KF\frac{\Vert K(t) - K^\star \Vert_F}{\Vert K(0) - K^\star \Vert_F} for the projected policy gradient initialized at several points.

Robustness to load variations is illustrated by demonstrating regulation performance under abrupt, piecewise-constant changes to the load admittance and reference values, comparing the synthesized optimal gain and several suboptimal alternatives.

Figure 4

Figure 3: Voltage v(t)v(t) trajectories for a time-varying load under LQI controllers synthesized with different gains and with online policy gradient adaptation.

Implications and Future Outlook

This work delivers practical and theoretical advances. The data-driven convex and policy-gradient-based LQI synthesis:

  • Enables direct controller design from input-state-output data with no model identification,
  • Bypasses the infeasibility of output feedback fixed-structure synthesis methods,
  • Affords robust tracking under reference changes and load disturbances,
  • Presents an avenue for online controller adaptation in evolving environments.

The simulation results substantiate the approach's utility for active distribution system control—a setting where load and grid parameters are seldom known a priori and evolve over time. Practical real-time tuning is facilitated by the policy gradient framework, which could be extended to incorporate stochasticity, estimation, and persistent excitation for closed-loop learning.

The theoretical implications extend to output regulation for nonlinear systems, bridging gaps between model-free reinforcement learning, convex control theory, and practical robust regulation. Challenges persist in guaranteeing robustness under process/model noise and in scaling to multi-agent or networked scenarios; these constitute promising avenues for subsequent research.

Conclusion

The paper develops and validates a rigorous data-driven methodology for synthesizing optimal LQI controllers for continuous-time linear systems using only empirical measurements. Both convex optimization and policy gradient-based synthesis are supported by formal guarantees and demonstrated numerically on a challenging power systems test case. The results furnish a scalable, theoretically substantiated toolkit for optimal output tracking in systems where physical models are only partially known or entirely unavailable.

(2604.14905)

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.