Papers
Topics
Authors
Recent
Search
2000 character limit reached

Centralized Gaussian Linear SCMs

Updated 11 January 2026
  • Centralized Gaussian Linear SCMs (CGL-SCMs) are Gaussian causal models where all exogenous variables are standardized to zero mean and unit variance, reducing parameter complexity.
  • They maintain full expressivity and observational equivalence to standard models, allowing for accurate identification and estimation of causal effects using graphical criteria.
  • An EM-based algorithm is employed for parameter learning, achieving high-fidelity causal effect estimation from finite sample data through closed-form interventions.

Centralized Gaussian Linear Structural Causal Models (CGL-SCMs) are a subclass of Gaussian Linear Structural Causal Models in which all exogenous variables (i.e., unobserved confounders and noises) are standardized to have zero mean and unit variance. This centralization eliminates the scale and location indeterminacy inherent in standard Gaussian Linear SCMs (GL-SCMs) by reducing the parameterization to a minimal yet fully expressive form. CGL-SCMs retain full expressivity with respect to observational and identifiable interventional distributions, enabling efficient parameter learning and causal effect estimation from finite samples using a specialized expectation–maximization (EM) procedure (Maiti et al., 8 Jan 2026).

1. Formal Specification and Expressivity

A Gaussian Linear SCM (GL-SCM) is defined by a tuple M′=⟨(U′,ε′),X,P,FX⟩M' = \langle (U', \varepsilon'), X, P, F_X\rangle, where U′∼N(μU′,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2) are multivariate normal confounders (with diagonal covariance), ε′∼N(με′,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2) are independent normal noise terms, and each endogenous variable XiX_i evolves via

Xi=∑j∈Pao(Xi)αjiXj+∑k∈Pau(Xi)αki′Uk′+μi′+εi′.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.

Edges Xj→XiX_j \to X_i and Uk′→XiU'_k \to X_i are present whenever αji≠0\alpha_{ji} \neq 0 and αki′≠0\alpha'_{ki} \neq 0.

A CGL-SCM is the special case where all exogenous variables are standardized: U∼N(0,I)U \sim \mathcal{N}(0, I), U′∼N(μU′,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)0, with endogenous variable structure

U′∼N(μU′,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)1

Centralization removes the means and variances of U′∼N(μU′,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)2, yielding a lower-dimensional parameter space.

Expressivity Theorem: For any GL-SCM U′∼N(μU′,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)3 with observed distribution U′∼N(μU′,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)4, there exists a CGL-SCM U′∼N(μU′,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)5 with the same graph U′∼N(μU′,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)6 such that U′∼N(μU′,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)7. Thus, CGL-SCMs and GL-SCMs are observationally indistinguishable and equally expressive in representing Gaussian-linear observational laws (Maiti et al., 8 Jan 2026).

2. Identifiability of Causal Effects

A query U′∼N(μU′,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)8 (e.g., U′∼N(μU′,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)9) is identifiable in a linear SCM with known graph ε′∼N(με′,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)0 if ε′∼N(με′,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)1 can be expressed uniquely in terms of the observational distribution ε′∼N(με′,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)2. Standard identification procedures such as Pearl's do-calculus and linear criteria (including instrument sets and graphical criteria of Brito–Pearl and Tian) extend directly to CGL-SCMs since these depend only on the topology ε′∼N(με′,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)3 and Gaussianity.

Identification Theorem: For a GL-SCM ε′∼N(με′,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)4 and corresponding CGL-SCM ε′∼N(με′,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)5 with ε′∼N(με′,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)6, identifiable queries ε′∼N(με′,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)7 satisfy ε′∼N(με′,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)8. This permits one to work always in the lower-dimensional, centralized parameterization without loss for identifiable causal effect estimation (Maiti et al., 8 Jan 2026).

An illustrative example: In the simple chain ε′∼N(με′,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)9 (no confounders), the CGL-SCM yields XiX_i0, and XiX_i1.

3. EM-Based Parameter Learning Algorithm

To estimate model parameters from data, the CGL-SCM admits a vectorized formulation. Let XiX_i2, XiX_i3, XiX_i4 the XiX_i5 weighted adjacency matrix of XiX_i6 (with XiX_i7), and XiX_i8 the length of the longest directed path. Define

XiX_i9

with Xi=∑j∈Pao(Xi)αjiXj+∑k∈Pau(Xi)αki′Uk′+μi′+εi′.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.0 the total sum of path weights from Xi=∑j∈Pao(Xi)αjiXj+∑k∈Pau(Xi)αki′Uk′+μi′+εi′.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.1 to Xi=∑j∈Pao(Xi)αjiXj+∑k∈Pau(Xi)αki′Uk′+μi′+εi′.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.2, Xi=∑j∈Pao(Xi)αjiXj+∑k∈Pau(Xi)αki′Uk′+μi′+εi′.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.3 the Xi=∑j∈Pao(Xi)αjiXj+∑k∈Pau(Xi)αki′Uk′+μi′+εi′.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.4 matrix of edges Xi=∑j∈Pao(Xi)αjiXj+∑k∈Pau(Xi)αki′Uk′+μi′+εi′.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.5, and Xi=∑j∈Pao(Xi)αjiXj+∑k∈Pau(Xi)αki′Uk′+μi′+εi′.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.6 intercepts. The stacked equations are

Xi=∑j∈Pao(Xi)αjiXj+∑k∈Pau(Xi)αki′Uk′+μi′+εi′.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.7

The joint Xi=∑j∈Pao(Xi)αjiXj+∑k∈Pau(Xi)αki′Uk′+μi′+εi′.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.8 is jointly Gaussian with explicitly computable mean and covariance.

EM Algorithm Steps:

  • E-step: For each data sample Xi=∑j∈Pao(Xi)αjiXj+∑k∈Pau(Xi)αki′Uk′+μi′+εi′.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.9, compute

Xj→XiX_j \to X_i0

Xj→XiX_j \to X_i1

  • M-step: Maximize the expected complete-data log-likelihood

Xj→XiX_j \to X_i2

with closed-form update for Xj→XiX_j \to X_i3:

Xj→XiX_j \to X_i4

Updates for Xj→XiX_j \to X_i5 and Xj→XiX_j \to X_i6 are performed by masked gradient ascent, preserving the zero-pattern dictated by the graph Xj→XiX_j \to X_i7.

EM guarantees non-decreasing observed-data likelihood at each iteration. Regularization (e.g., Xj→XiX_j \to X_i8-penalties) and early stopping are advisable for small Xj→XiX_j \to X_i9 to prevent overfitting (Maiti et al., 8 Jan 2026).

4. Causal Inference and Effect Estimation

After model fitting, causal queries are evaluated by modifying structural equations and computing the resulting Gaussian distribution, as dictated by do-calculus.

Do-Interventions: For intervention Uk′→XiU'_k \to X_i0, incoming edges to Uk′→XiU'_k \to X_i1 are removed (i.e., zeroed in Uk′→XiU'_k \to X_i2), and Uk′→XiU'_k \to X_i3 is set to Uk′→XiU'_k \to X_i4. Remaining Uk′→XiU'_k \to X_i5 are solved as linear functions of Uk′→XiU'_k \to X_i6 and Uk′→XiU'_k \to X_i7. The post-interventional distribution Uk′→XiU'_k \to X_i8 remains multivariate normal, with parameters derived from the submatrices of the modified Uk′→XiU'_k \to X_i9 and αji≠0\alpha_{ji} \neq 00.

Example (Linear Chain):

Chain Structure Total Effect αji≠0\alpha_{ji} \neq 01
αji≠0\alpha_{ji} \neq 02 αji≠0\alpha_{ji} \neq 03 αji≠0\alpha_{ji} \neq 04Yαji≠0\alpha_{ji} \neq 05

This pipeline applies to any graph-identifiable query, including counterfactuals, due to the closed-form propagation properties of Gaussian-linear models (Maiti et al., 8 Jan 2026).

5. Empirical Evaluation and Application

Synthetic validation was conducted using the "frontdoor" and "napkin" benchmark graphs:

  • Frontdoor graph: (three observed nodes αji≠0\alpha_{ji} \neq 06 with unobserved confounder αji≠0\alpha_{ji} \neq 07)
  • Napkin graph: (four observed nodes, two latent confounders)
  • In both cases, αji≠0\alpha_{ji} \neq 08 samples were drawn from known CGL-SCMs.

After learning with the EM algorithm, the estimated causal effects closely matched ground truth. In the frontdoor scenario:

  • True αji≠0\alpha_{ji} \neq 09, estimated as αki′≠0\alpha'_{ki} \neq 00.

For the napkin graph:

  • True αki′≠0\alpha'_{ki} \neq 01, estimated as αki′≠0\alpha'_{ki} \neq 02.

Mean and variance estimates were consistently within a few percent of their true values, demonstrating high-fidelity recovery of causal effects from finite-sample observational data using the CGL-SCM EM algorithm (Maiti et al., 8 Jan 2026).

6. Parameter Reduction and Practical Advantages

CGL-SCMs achieve parameter reduction by standardizing all exogenous variables (confounders and noises) to zero mean and unit variance. This eliminates the latent scaling and location degrees of freedom in general GL-SCMs: the means and variances of αki′≠0\alpha'_{ki} \neq 03 are removed from the model specification. As a result, the number of free parameters—particularly those associated with unobserved confounders—is drastically reduced. Despite this, the class retains full expressivity over both observational and graph-identifiable interventional distributions. This simplification is particularly advantageous for finite-sample learning, where overparameterization often leads to infeasible or unstable estimation in the presence of unobserved confounding (Maiti et al., 8 Jan 2026).

The EM-based learning algorithm accommodates this streamlined parameterization and enables efficient estimation of edge-weights and bias terms, ensuring that causal queries remain representable and computable in closed form after training.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Centralized Gaussian Linear SCMs (CGL-SCMs).