Papers
Topics
Authors
Recent
Search
2000 character limit reached

Centralized Gaussian Linear SCMs

Updated 11 January 2026
  • Centralized Gaussian Linear SCMs (CGL-SCMs) are Gaussian causal models where all exogenous variables are standardized to zero mean and unit variance, reducing parameter complexity.
  • They maintain full expressivity and observational equivalence to standard models, allowing for accurate identification and estimation of causal effects using graphical criteria.
  • An EM-based algorithm is employed for parameter learning, achieving high-fidelity causal effect estimation from finite sample data through closed-form interventions.

Centralized Gaussian Linear Structural Causal Models (CGL-SCMs) are a subclass of Gaussian Linear Structural Causal Models in which all exogenous variables (i.e., unobserved confounders and noises) are standardized to have zero mean and unit variance. This centralization eliminates the scale and location indeterminacy inherent in standard Gaussian Linear SCMs (GL-SCMs) by reducing the parameterization to a minimal yet fully expressive form. CGL-SCMs retain full expressivity with respect to observational and identifiable interventional distributions, enabling efficient parameter learning and causal effect estimation from finite samples using a specialized expectation–maximization (EM) procedure (Maiti et al., 8 Jan 2026).

1. Formal Specification and Expressivity

A Gaussian Linear SCM (GL-SCM) is defined by a tuple M=(U,ε),X,P,FXM' = \langle (U', \varepsilon'), X, P, F_X\rangle, where UN(μU,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2) are multivariate normal confounders (with diagonal covariance), εN(με,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2) are independent normal noise terms, and each endogenous variable XiX_i evolves via

Xi=jPao(Xi)αjiXj+kPau(Xi)αkiUk+μi+εi.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.

Edges XjXiX_j \to X_i and UkXiU'_k \to X_i are present whenever αji0\alpha_{ji} \neq 0 and αki0\alpha'_{ki} \neq 0.

A CGL-SCM is the special case where all exogenous variables are standardized: UN(0,I)U \sim \mathcal{N}(0, I), UN(μU,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)0, with endogenous variable structure

UN(μU,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)1

Centralization removes the means and variances of UN(μU,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)2, yielding a lower-dimensional parameter space.

Expressivity Theorem: For any GL-SCM UN(μU,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)3 with observed distribution UN(μU,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)4, there exists a CGL-SCM UN(μU,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)5 with the same graph UN(μU,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)6 such that UN(μU,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)7. Thus, CGL-SCMs and GL-SCMs are observationally indistinguishable and equally expressive in representing Gaussian-linear observational laws (Maiti et al., 8 Jan 2026).

2. Identifiability of Causal Effects

A query UN(μU,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)8 (e.g., UN(μU,Σ2)U' \sim \mathcal{N}(\mu_{U'}, \Sigma^2)9) is identifiable in a linear SCM with known graph εN(με,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)0 if εN(με,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)1 can be expressed uniquely in terms of the observational distribution εN(με,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)2. Standard identification procedures such as Pearl's do-calculus and linear criteria (including instrument sets and graphical criteria of Brito–Pearl and Tian) extend directly to CGL-SCMs since these depend only on the topology εN(με,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)3 and Gaussianity.

Identification Theorem: For a GL-SCM εN(με,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)4 and corresponding CGL-SCM εN(με,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)5 with εN(με,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)6, identifiable queries εN(με,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)7 satisfy εN(με,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)8. This permits one to work always in the lower-dimensional, centralized parameterization without loss for identifiable causal effect estimation (Maiti et al., 8 Jan 2026).

An illustrative example: In the simple chain εN(με,Ψ2)\varepsilon' \sim \mathcal{N}(\mu_{\varepsilon'}, \Psi^2)9 (no confounders), the CGL-SCM yields XiX_i0, and XiX_i1.

3. EM-Based Parameter Learning Algorithm

To estimate model parameters from data, the CGL-SCM admits a vectorized formulation. Let XiX_i2, XiX_i3, XiX_i4 the XiX_i5 weighted adjacency matrix of XiX_i6 (with XiX_i7), and XiX_i8 the length of the longest directed path. Define

XiX_i9

with Xi=jPao(Xi)αjiXj+kPau(Xi)αkiUk+μi+εi.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.0 the total sum of path weights from Xi=jPao(Xi)αjiXj+kPau(Xi)αkiUk+μi+εi.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.1 to Xi=jPao(Xi)αjiXj+kPau(Xi)αkiUk+μi+εi.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.2, Xi=jPao(Xi)αjiXj+kPau(Xi)αkiUk+μi+εi.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.3 the Xi=jPao(Xi)αjiXj+kPau(Xi)αkiUk+μi+εi.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.4 matrix of edges Xi=jPao(Xi)αjiXj+kPau(Xi)αkiUk+μi+εi.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.5, and Xi=jPao(Xi)αjiXj+kPau(Xi)αkiUk+μi+εi.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.6 intercepts. The stacked equations are

Xi=jPao(Xi)αjiXj+kPau(Xi)αkiUk+μi+εi.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.7

The joint Xi=jPao(Xi)αjiXj+kPau(Xi)αkiUk+μi+εi.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.8 is jointly Gaussian with explicitly computable mean and covariance.

EM Algorithm Steps:

  • E-step: For each data sample Xi=jPao(Xi)αjiXj+kPau(Xi)αkiUk+μi+εi.X_i = \sum_{j\in\mathrm{Pa}^o(X_i)} \alpha_{j i} X_j + \sum_{k\in \mathrm{Pa}^u(X_i)} \alpha'_{k i} U'_k + \mu'_i + \varepsilon'_i.9, compute

XjXiX_j \to X_i0

XjXiX_j \to X_i1

  • M-step: Maximize the expected complete-data log-likelihood

XjXiX_j \to X_i2

with closed-form update for XjXiX_j \to X_i3:

XjXiX_j \to X_i4

Updates for XjXiX_j \to X_i5 and XjXiX_j \to X_i6 are performed by masked gradient ascent, preserving the zero-pattern dictated by the graph XjXiX_j \to X_i7.

EM guarantees non-decreasing observed-data likelihood at each iteration. Regularization (e.g., XjXiX_j \to X_i8-penalties) and early stopping are advisable for small XjXiX_j \to X_i9 to prevent overfitting (Maiti et al., 8 Jan 2026).

4. Causal Inference and Effect Estimation

After model fitting, causal queries are evaluated by modifying structural equations and computing the resulting Gaussian distribution, as dictated by do-calculus.

Do-Interventions: For intervention UkXiU'_k \to X_i0, incoming edges to UkXiU'_k \to X_i1 are removed (i.e., zeroed in UkXiU'_k \to X_i2), and UkXiU'_k \to X_i3 is set to UkXiU'_k \to X_i4. Remaining UkXiU'_k \to X_i5 are solved as linear functions of UkXiU'_k \to X_i6 and UkXiU'_k \to X_i7. The post-interventional distribution UkXiU'_k \to X_i8 remains multivariate normal, with parameters derived from the submatrices of the modified UkXiU'_k \to X_i9 and αji0\alpha_{ji} \neq 00.

Example (Linear Chain):

Chain Structure Total Effect αji0\alpha_{ji} \neq 01
αji0\alpha_{ji} \neq 02 αji0\alpha_{ji} \neq 03 αji0\alpha_{ji} \neq 04Yαji0\alpha_{ji} \neq 05

This pipeline applies to any graph-identifiable query, including counterfactuals, due to the closed-form propagation properties of Gaussian-linear models (Maiti et al., 8 Jan 2026).

5. Empirical Evaluation and Application

Synthetic validation was conducted using the "frontdoor" and "napkin" benchmark graphs:

  • Frontdoor graph: (three observed nodes αji0\alpha_{ji} \neq 06 with unobserved confounder αji0\alpha_{ji} \neq 07)
  • Napkin graph: (four observed nodes, two latent confounders)
  • In both cases, αji0\alpha_{ji} \neq 08 samples were drawn from known CGL-SCMs.

After learning with the EM algorithm, the estimated causal effects closely matched ground truth. In the frontdoor scenario:

  • True αji0\alpha_{ji} \neq 09, estimated as αki0\alpha'_{ki} \neq 00.

For the napkin graph:

  • True αki0\alpha'_{ki} \neq 01, estimated as αki0\alpha'_{ki} \neq 02.

Mean and variance estimates were consistently within a few percent of their true values, demonstrating high-fidelity recovery of causal effects from finite-sample observational data using the CGL-SCM EM algorithm (Maiti et al., 8 Jan 2026).

6. Parameter Reduction and Practical Advantages

CGL-SCMs achieve parameter reduction by standardizing all exogenous variables (confounders and noises) to zero mean and unit variance. This eliminates the latent scaling and location degrees of freedom in general GL-SCMs: the means and variances of αki0\alpha'_{ki} \neq 03 are removed from the model specification. As a result, the number of free parameters—particularly those associated with unobserved confounders—is drastically reduced. Despite this, the class retains full expressivity over both observational and graph-identifiable interventional distributions. This simplification is particularly advantageous for finite-sample learning, where overparameterization often leads to infeasible or unstable estimation in the presence of unobserved confounding (Maiti et al., 8 Jan 2026).

The EM-based learning algorithm accommodates this streamlined parameterization and enables efficient estimation of edge-weights and bias terms, ensuring that causal queries remain representable and computable in closed form after training.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Centralized Gaussian Linear SCMs (CGL-SCMs).