---
title: Acyclic Graphical Continuous Lyapunov Models
url: https://www.emergentmind.com/topics/acyclic-graphical-continuous-lyapunov-models-gclms
type: topic
---

# Acyclic Graphical Continuous Lyapunov Models

Searching arXiv for recent papers on acyclic graphical continuous Lyapunov models and related identifiability results.
Acyclic graphical continuous Lyapunov models (GCLMs) are a subclass of graphical continuous Lyapunov models in which the directed support of the drift matrix is a directed acyclic graph (DAG). They model each observation as a cross-sectional draw from the stationary law of a stable continuous-time linear stochastic system, so that dependence is encoded through a continuous Lyapunov equation rather than through a structural equation solved at a single time point. In the Gaussian formulation, the stationary distribution is determined by the drift matrix and the noise covariance; in the non-Gaussian formulation, higher-order cumulants of the stationary law satisfy generalized Lyapunov equations and can identify edge weights beyond covariance information alone [2005.10483; 2603.17142].

## 1. Stochastic formulation and graph semantics

In the Gaussian setting, a GCLM starts from a multivariate Ornstein–Uhlenbeck process
$$
\mathrm dX(t)=M\,(X(t)-a)\,\mathrm dt + D\,\mathrm dW(t),
$$
with drift matrix \(M\in\mathbb R^{p\times p}\), volatility matrix \(D\in\mathbb R^{p\times p}\), and \(C=D D^\top\). If \(M\) is Hurwitz, then \(X(t)\) has a unique stationary Gaussian law \(N(a,\Sigma)\), and \(\Sigma\) is the unique positive-definite solution of
$$
M\Sigma+\Sigma M^\top + C = 0.
$$
A graphical continuous Lyapunov model imposes a zero pattern on \(M\) through a directed graph \(G=(V,E)\), with an edge \(i\to j\) indicating that \(m_{ji}\) may be nonzero; self-loops \(i\to i\) are always included so that diagonal entries remain free [2209.03835].

For fixed \(C\succ 0\), the model can be written as
$$
\mathcal M_{G,C}:=\{\Sigma\in PD_p:\exists\,M\in \mathrm{Stab}(E)\text{ with }M\Sigma+\Sigma M^\top=-C\},
$$
where \(\mathrm{Stab}(E)\subseteq \mathbb R^E\) denotes the stable matrices with the prescribed sparsity pattern [2209.03835]. In the acyclic case, one may topologically order the nodes so that \(i\to j\Rightarrow i\le j\), making the off-diagonal support triangular after reordering; stability is then typically enforced through negative diagonal entries [2005.10483].

The broader non-Gaussian extension replaces Brownian noise by a Lévy process. In the formulation studied by Recke and Hansen, one observes a cross-sectional sample \(X\) from the stationary solution of
$$
\mathrm dX_t=A X_t\,\mathrm dt+\mathrm dZ_t,
$$
where \(A\in\mathbb R^{d\times d}\) is stable, \(A_{ij}=0\) whenever \(j\to i\notin E\), and \(Z_t\) is a \(d\)-dimensional Lévy process. Under mild integrability theorems, the stationary law exists uniquely and admits the representation
$$
X\overset d= \int_0^\infty e^{As}\,\mathrm dZ_s.
$$
This places acyclic GCLMs inside a larger class of stationary linear Markov-process models whose observational content is entirely cross-sectional [2603.17142].

## 2. Lyapunov equations, trek expansions, and covariance geometry

The defining covariance relation is the continuous Lyapunov equation
$$
A\Sigma+\Sigma A^\top+\Omega=0,
$$
or, after vectorization,
$$
(I\otimes A + A\otimes I)\,\mathrm{vec}(\Sigma)+\mathrm{vec}(\Omega)=0.
$$
This vectorized form is central both for identifiability arguments and for estimation procedures, because it turns the covariance constraints into linear equations in \(\mathrm{vec}(A)\) once \(\Sigma\) is regarded as fixed [2407.21223].

Hansen’s trek rule gives a graphical expansion of \(\Sigma_{ij}\) in terms of the drift and volatility parameters. In the general mixed-graph setting, if \(\Lambda=A+I\) has spectral radius \(<1\), then
$$
\Sigma_{ij}
=
\sum_{\tau\in\mathcal T(i,j)}
2^{-\,l(\tau)-1}\binom{l(\tau)}{n(\tau)}\,
\omega(\Lambda,\Omega,\tau),
$$
where \(\tau\) ranges over treks, \(n(\tau)\) and \(m(\tau)\) count the left and right directed parts, \(l(\tau)=n(\tau)+m(\tau)\), and \(\omega(\cdot)\) is the trek weight [2407.21223]. In the acyclic case, after rescaling so that \(A_{ii}=-1\), the expansion becomes a finite polynomial over treks in the acyclic base graph:
$$
\Sigma_{ij}
=
\sum_{\tau\in\mathcal T_0(i,j)}
2^{-\,l(\tau)-1}\binom{l(\tau)}{n(\tau)}\,
\omega(A,\Omega,\tau).
$$
Thus, for acyclic GCLMs, each covariance entry is a polynomial in the off-diagonal drift parameters and the volatility entries rather than an infinite series [2407.21223].

The trek perspective clarifies how acyclic Lyapunov models differ from linear additive-noise DAG models. The Lyapunov trek rule carries the extra combinatorial factors \(2^{-l-1}\binom{l}{n}\), and even in DAG cases the induced covariance geometry is not the same as in algebraic structural-equation models. In particular, the star-graph formulas in Hansen’s analysis show that the implied \(\Sigma\) need not satisfy the rank-one tetrad constraints of classical factor models [2407.21223].

For special acyclic drifts, explicit recursions are available. In the path-graph case with \(A_{ii}=-1\), \(A_{i,i-1}=\zeta\), and \(\Omega=\gamma I\), the covariances satisfy
$$
\Sigma_{ij}
=
\frac{\zeta}{2}\Sigma_{i-1,j}
+
\frac{\zeta}{2}\Sigma_{i,j-1}
+
\frac{\gamma}{2}\delta_{ij}.
$$
For certain sign patterns and diagonal \(\Omega\), Hansen also derives the lower bound
$$
\Sigma_{dd}\ge -\frac{\Omega_{dd}}{2m_{dd}}+\frac12 A_{d1}\Sigma_{11}A_{d1}^\top,
$$
and in simple path models the marginal variances strictly increase along the topological order [2407.21223]. This suggests that acyclicity constrains not only identifiability but also qualitative variance propagation.

## 3. Identifiability from covariance in the Gaussian acyclic case

In the Gaussian setting with fixed \(C\succ 0\), identifiability asks whether the map
$$
\phi_{G,C}: \mathrm{Stab}(E)\to PD_p,\qquad M\mapsto \Sigma,
$$
is injective. The fiber at \(M_0\) is
$$
\mathcal F_{G,C}(M_0)
=
\{M\in\mathrm{Stab}(E):\phi_{G,C}(M)=\phi_{G,C}(M_0)\}.
$$
The model is globally identifiable if every fiber is a singleton, generically identifiable if singleton fibers fail only on an algebraic exception set, and non-identifiable if all fibers are infinite [2209.03835].

Dettling, Homs, Améndola, Drton, and Hansen prove the basic characterization: for arbitrary directed graphs with self-loops and fixed \(C\succ 0\), the parametrization is globally identifiable if and only if the graph has no directed \(2\)-cycle. Since every DAG is simple in this sense, every acyclic GCLM is globally identifiable [2209.03835]. In the diagonal-noise case, the same equivalence holds as a corollary.

The proof relies on vectorized rank conditions. Keeping only the symmetric equations yields a \(p(p+1)/2\times p^2\) matrix \(A(\Sigma)\) such that
$$
A(\Sigma)\,\mathrm{vec}(M)=-\mathrm{vech}(C).
$$
If \(A(\Sigma)_{\cdot,E}\) denotes the submatrix formed by the columns corresponding to the allowed edges, then
\[
\mathcal M_{G,C}\text{ is globally identifiable } \Leftrightarrow
\forall\,\Sigma\in\mathcal M_{G,C},\ \mathrm{rank}\,A(\Sigma)_{\cdot,E}=|E|,
\]
and
\[
\mathcal M_{G,C}\text{ is generically identifiable } \Leftrightarrow
\exists\,\Sigma\in\mathcal M_{G,C}\text{ with }\mathrm{rank}\,A(\Sigma)_{\cdot,E}=|E|.
\]
For DAGs, a topological ordering makes \(A(\Sigma)_{\cdot,E}\) block-upper-triangular with positive determinants on the diagonal blocks, which yields full column rank for all \(\Sigma\) in the model [2209.03835].

This covariance-based result distinguishes acyclic GCLMs from non-simple graphs, where directed \(2\)-cycles create immediate non-uniqueness. It also places acyclic GCLMs closer to identifiable sparse covariance parametrizations than to generic cyclic Lyapunov models. A frequent misconception is that covariance alone is intrinsically insufficient to identify direction in continuous-time linear systems; within the Gaussian GCLM framework that statement is false for DAG supports with known \(C\), because acyclicity already guarantees global identifiability [2209.03835].

## 4. Higher-order cumulants and non-Gaussian acyclic models

The non-Gaussian theory extends the second-order Lyapunov equation to cumulant tensors of arbitrary order. If \(C_k=\mathrm{cum}_k(Z_1)\) is the \(k\)-th cumulant tensor of the driving Lévy noise and \(K=\mathrm{cum}_k(X)\) is the \(k\)-th cumulant tensor of the stationary law, then \(K\) satisfies
$$
K\times_1 A + K\times_2 A + \cdots + K\times_k A + C_k = 0.
$$
Equivalently, in Tucker notation,
$$
\sum_{n=1}^k (I\otimes\cdots\otimes A\otimes\cdots\otimes I)\,\mathrm{vec}(K)+\mathrm{vec}(C_k)=0.
$$
For \(k=2\) this reduces to the classical covariance Lyapunov equation; for \(k=3\) one obtains the componentwise identity
$$
A_{i_1\alpha}K_{\alpha i_2 i_3}
+
A_{i_2\beta}K_{i_1\beta i_3}
+
A_{i_3\gamma}K_{i_1 i_2 \gamma}
+
(C_3)_{i_1 i_2 i_3}
=0
$$
[2603.17142].

The identifiability problem changes fundamentally when the noise is Gaussian. If \(C_3\equiv 0\), then \(A\) often cannot be recovered from \(\Sigma\) alone in the unrestricted continuous model. Recke and Hansen therefore assume nonzero higher-order coordinate cumulants, as occurs for independent non-Gaussian jumps with diagonal \(C_k\). Their main theorem states that if \(G\) is a connected DAG with self-loops and \(r\ge 3\), then outside a proper algebraic exception set the triple \((A,C_2,C_r)\) is determined uniquely up to a common scalar factor by \((\Sigma,K^{(r)})\). Equivalently,
$$
\phi_r:(A,C_2,C_r)\mapsto (\Sigma,K^{(r)})
$$
has \(1\)-dimensional generic fiber
$$
\{c\,(A,C_2,C_r):c>0\}.
$$
The residual ambiguity is the trivial time-rescaling \(A\mapsto cA\), \(C_k\mapsto cC_k\) [2603.17142].

The proof strategy vectorizes the second- and \(r\)-th-order equations simultaneously, removes the rows and columns associated with diagonal cumulant parameters, and reduces the problem to a linear system in \(\mathrm{vec}(A)\),
$$
\begin{bmatrix}
A_2(\Sigma)_{\mathrm{off}}\\
A_r(K)_{\mathrm{off}}
\end{bmatrix}
\mathrm{vec}(A)=0.
$$
A special stable lower-triangular choice of \(A\) together with generic diagonal cumulants yields rank \(d^2-1\), and algebraic-geometric arguments then show that the one-dimensional kernel is generic [2603.17142].

A plausible implication is that acyclicity alone does not exhaust the identifiability structure of continuous Lyapunov models: once higher-order cumulants are observed, the same DAG restriction supports identification in broader non-Gaussian equilibrium models than the Gaussian covariance theory initially suggests.

## 5. Model equivalence and structural identifiability of DAGs

When the graph itself is unknown, a distinct question arises: which DAGs define the same acyclic GCLM? In the setting of stable drift matrices with uncorrelated noise
$$
\Omega=\mathrm{diag}(\omega_1^2,\dots,\omega_n^2),
$$
the acyclic GCLM of \(G\) is
$$
M_G
=
\{\Sigma\succ 0:\exists\text{ stable }A\text{ with sparsity }G\text{ solving }A\Sigma+\Sigma A^\top+\Omega=0\},
$$
and two DAGs are model equivalent if \(M_{G_1}=M_{G_2}\) [2510.04985].

The first characterization is the \(4\)-node criterion: two DAGs \(G_1,G_2\) on the same vertex set are model equivalent if and only if they have the same skeleton and, for every \(4\)-element subset \(K\subseteq V\), the induced subgraphs \(G_1[K]\) and \(G_2[K]\) are themselves model equivalent [2510.04985]. This is the Lyapunov analogue of the Verma–Pearl characterization for Bayesian networks, but the local obstruction size increases from \(3\) nodes to \(4\).

The second characterization is transformational. An edge \(i\to j\) is super-covered if

- \(\mathrm{pa}(i)\cup\{i\}=\mathrm{pa}(j)\) and \(\mathrm{ch}(i)=\mathrm{ch}(j)\cup\{j\}\),
- every parent of \(j\) is a parent of every child of \(i\),
- any third vertex \(k\notin\{i,j\}\) is either a neighbor of both \(i\) and \(j\) or else has no trek connecting it to \(i\) or to \(j\).

Two simple DAGs satisfy \(M_{G_1}=M_{G_2}\) if and only if one can transform \(G_1\) into \(G_2\) by a finite sequence of super-covered edge flips, and if they differ on \(\delta\) edges then exactly \(\delta\) flips suffice [2510.04985].

These results have two important consequences. First, Lyapunov equivalence classes refine Bayesian-network equivalence classes: every pair of Lyapunov-equivalent DAGs is also Markov-equivalent as Bayesian networks, but not conversely. Second, both equivalence testing and structural identifiability admit polynomial-time algorithms. The equivalence test runs in \(O(n^4)\) by checking skeleton equality and all induced \(4\)-node subgraphs; structural identifiability can be tested in \(O(n^3)\) by searching for super-covered edges via the seven non-identifiable \(4\)-node patterns [2510.04985]. This corrects the common assumption that continuous-time DAG models inherit exactly the same equivalence theory as Gaussian Bayesian networks.

## 6. Estimation, model selection, and finite-sample behavior

Estimation methods for acyclic GCLMs split naturally into covariance-based Gaussian procedures and cumulant-based non-Gaussian procedures. In the non-Gaussian setting, Recke and Hansen propose a semiparametric moment-matching estimator. Given \(n\) i.i.d. samples \(X^{(1)},\dots,X^{(n)}\), one computes empirical \(\widehat\Sigma\) and \(\widehat K^{(r)}\), forms the estimated coefficient matrix
$$
\widehat A=
\begin{bmatrix}
A_2(\widehat\Sigma)_{\mathrm{off}}\\
A_r(\widehat K)_{\mathrm{off}}
\end{bmatrix},
$$
and estimates \(A\) by the right singular vector corresponding to the smallest singular value:
$$
\widehat A = \mathrm{vec}^{-1}(v_{\min}(\widehat A)),
$$
normalized by \(\|\widehat A\|_F=1\) and the sign convention \(\mathrm{tr}(\widehat A)<0\). Under finite \(2r\)-th moments,
$$
\sqrt n\,\mathrm{vec}(\widehat A-A)\to
\mathcal N\!\left(
0,\,
(\mathrm{vec}(A)^\top\otimes A^+)\,
\mathcal M\,
(\mathrm{vec}(A)\otimes (A^+)^\top)
\right),
$$
so the estimator is \(\sqrt n\)-consistent and asymptotically normal [2603.17142].

The same paper emphasizes the difficulty of the unconstrained estimation problem in finite samples. In simulations with \(d\) up to approximately \(12\) and compound-Poisson noise, the estimator showed noticeable bias unless \(n\gg 1000\); the root-\(n\)-scaled bias decayed slowly when the noise correlation \(\rho\) was large. To achieve \(\mathrm{MSE}\lesssim 10^{-2}\), sample sizes of order \(10^4\)–\(10^5\) were often needed, with the requirement increasing in both dimension and correlation [2603.17142]. Thus higher-order identifiability does not by itself imply easy estimation.

For Gaussian sparse learning, the Direct Lyapunov Lasso solves
$$
\widehat M
=
\arg\min_{M\in\mathbb R^{p\times p}}
\left\{
\frac12\|M\widehat\Sigma+\widehat\Sigma M^\top + C\|_F^2
+\lambda\|M\|_1
\right\},
$$
or, after vectorization, a standard lasso with Hessian \(\Gamma(\widehat\Sigma)=A(\widehat\Sigma)^\top A(\widehat\Sigma)\). Under the irrepresentability condition
$$
\left\|\Gamma^*_{S^cS}(\Gamma^*_{SS})^{-1}\right\|_\infty<1-\alpha,
$$
together with explicit sample-size and tuning-parameter inequalities, one obtains uniqueness, sup-norm error control, and exact support and sign recovery when the minimum signal exceeds the error bound [2208.13572]. The analysis also shows that this condition is unexpectedly hard to satisfy.

Acyclicity is central here as well. For DAG supports, if one starts at a diagonal drift \(M^0=\mathrm{diag}(-d_1,\dots,-d_p)\), then the irrepresentability condition holds uniformly in a neighborhood of \(M^0\) precisely when
$$
d_i<d_j\qquad\text{whenever } i\to j \text{ is an edge of }G,
$$
an ordering that is possible only for acyclic graphs [2208.13572]. In contrast, simulations on cyclic supports showed that the condition almost never holds. Nonetheless, numerical experiments found that the lasso recovered relevant structure robustly even under volatility misspecification, and on the Sachs et al. protein-signaling data a selected graph on dataset \(7\) had \(17\) directed edges, with roughly \(80\)–\(90\%\) matching or reversing edges in the published network [2208.13572].

The original structure-learning approach of Varando and Hansen instead minimizes a nonconvex penalized loss in \((B,C)\),
$$
\min_{B\ \mathrm{stable},\, C\ \mathrm{diagonal}}
L(\Sigma(B,C))
+\lambda\sum_{i\neq j}|B_{ij}|
+\kappa\|C-I_p\|_F^2,
$$
with either Gaussian negative log-likelihood or Frobenius loss, and uses a proximal-gradient algorithm in which each gradient step is computed through Lyapunov solves of cost \(\mathcal O(p^3)\) [2005.10483]. In historical terms, this optimization framework introduced sparse learning for GCLMs, while later work clarified when acyclicity yields exact covariance identifiability, when model equivalence remains nontrivial, and when higher-order cumulants can recover drift parameters beyond the Gaussian case.

Acyclic GCLMs therefore occupy a technically distinctive position among continuous-time graphical models: they are globally identifiable from covariance under fixed Gaussian noise, admit a refined equivalence theory that is stricter than Bayesian-network Markov equivalence, and remain identifiable in broader non-Gaussian settings through higher-order cumulants, while still presenting substantial finite-sample and model-selection challenges [2209.03835; 2510.04985; 2603.17142].

Source: https://www.emergentmind.com/topics/acyclic-graphical-continuous-lyapunov-models-gclms