---
title: Proximal Linear Structural Equations Model
url: https://www.emergentmind.com/topics/proximal-linear-structural-equations-model
type: topic
---

# Proximal Linear Structural Equations Model

A proximal linear structural equations model is a linear SEM used to analyze causal effects in the presence of hidden confounding by exploiting observed proxies, typically treatment-inducing proxies \(Z\) and outcome-inducing proxies \(W\). In the proximal causal inference literature, such models supply bridge equations, moment restrictions, and estimators that recover the causal effect of a treatment \(A\) or \(D\) on an outcome \(Y\) even when an unobserved confounder \(U\) is present, provided that proxy validity, completeness, and relevance conditions hold or are replaced by structured restrictions on proxy invalidity [2208.00105; 2507.19623]. In a separate SEM literature, the term “proximal” refers instead to proximal optimization algorithms for convex relaxations of regularized path-analysis models, not to proxy-based identification [1809.06156].

## 1. Canonical model and notation

In proximal causal inference, the observed data may be written as i.i.d. draws
\[
D_i=(A_i,Y_i,Z_i,W_i,X_i),
\]
where \(U\) is an unobserved confounder, \(X\) denotes observed covariates, \(A\in\{0,1\}\) or \(D\in\mathbb R\) is the treatment, \(Z\) is a negative-control exposure proxy or treatment-inducing proxy, \(W\) is a negative-control outcome proxy or outcome-inducing proxy, and \(Y\) is the outcome [2208.00105; 2507.19623].

A mathematically explicit LSEM used for bias analysis specifies that \(A\) follows an \(\mathrm{expit}\) model in \((U,X)\), that \(Z\) and \(W\) depend linearly on \(U\) and \(X\), and that the potential outcome \(Y(a)\) depends linearly on \(a\), \(X\), \(U\), and the interaction \(aU\) [2208.00105]. This formulation is designed to study what occurs when proximal identification assumptions fail.

A canonical proximal linear SEM used for identification with potentially invalid treatment proxies is
\[
E[Y_i\mid D_i,Z_i,U_i] = \beta D_i + \alpha^T Z_i + \beta_u U_i,
\]
\[
E[W_i\mid D_i,Z_i,U_i] = \eta_u U_i.
\]
Here \(\beta\) is the causal effect of \(D\) on \(Y\), \(\alpha=(\alpha_1,\dots,\alpha_{p_z})^T\) captures direct effects of the \(Z\)-proxies on \(Y\), and \(\beta_u,\eta_u\) are nuisance loadings of \(U\) [2507.19623]. From the second equation,
\[
U_i = E[W_i\mid D_i,Z_i]/\eta_u,
\]
so the model implies
\[
E[Y_i\mid D_i,Z_i] = \beta D_i + \alpha^T Z_i + \gamma\,E[W_i\mid D_i,Z_i],
\qquad \gamma=\beta_u/\eta_u.
\]
This leads to the observed regression form
\[
Y_i = \beta D_i + Z_i^T\alpha + W_i\gamma + \varepsilon_i,
\qquad E[\varepsilon_i\mid D_i,Z_i]=0,
\]
with \(\operatorname{Var}(\varepsilon_i\mid D_i,Z_i)=\sigma_\varepsilon^2\) [2507.19623].

Within this formulation, a \(Z\)-proxy \(j\) is valid iff \(\alpha_j=0\); otherwise it is invalid. \(W\) is a valid outcome-inducing proxy because its conditional mean does not depend on \(D\) or \(Z\) beyond \(U\) [2507.19623].

## 2. Identification via proxy restrictions and bridge functions

The proximal identification program rests on a set of assumptions that extend standard causal conditions by introducing proxy restrictions [2208.00105]. Beyond consistency and positivity, the key assumptions are negative-control proxy validity,
\[
Y(a,z)=Y(a)\quad\forall a,z,\qquad W(a,z)=W\quad\forall a,z,
\]
latent unconfoundedness,
\[
(Z\perp (Y(a),W)\mid U,X), \qquad (W\perp A\mid U,X),
\]
completeness, and \(U\)-relevance [2208.00105].

Under these assumptions, there exists an outcome-bridge \(h(w,a,x)\) satisfying
\[
E[Y\mid Z,A,X]=\int h(w,A,X)\,dF(w\mid Z,A,X),
\]
which yields the proximal g-formula
\[
E[Y(a)] = E[h(W,a,X)].
\]
This bridge representation is central because it translates latent-confounding adjustment into an estimable functional equation involving observables [2208.00105].

In the canonical linear SEM with many treatment-inducing proxies, identification can be expressed through population moments. Let \(M=[Z\mid D]\). Then the model imposes
\[
E\!\left[M^T(Y-D\beta-Z\alpha-W\gamma)\right]=0.
\]
Define
\[
\Gamma^*=(E[M^TM])^{-1}E[M^TY],\qquad
\delta^*=(E[M^TM])^{-1}E[M^TW].
\]
Partition \(\Gamma^*=(\Gamma_Z^*,\Gamma_D^*)\) and similarly \(\delta^*\). The moment equations imply
\[
\Gamma_Z^*=\alpha+\delta_Z^*\gamma,\qquad
\Gamma_D^*=\beta+\delta_D^*\gamma.
\]
If \(A=\{j:\alpha_j\neq 0\}\) is the set of invalid \(Z\)-proxies and \(s_z=|A|\), Theorem 1 states that there is a unique solution \((\alpha,\gamma)\) iff for every subset \(C\subseteq\{1,\dots,p_z\}\) with \(|C|=p_z-I+1\), there exists a constant \(q_C\) such that \(\delta_j^* q_C=\Gamma_j^*\) for all \(j\in C\), and \(q_C\) must be the same for all such \(C\) [2507.19623]. A simple corollary is the majority rule: if \(I\le p_z/2\), meaning fewer than half of the \(Z\)-proxies are invalid, the solution is automatically unique. Once \((\alpha,\gamma)\) is identified,
\[
\beta=\Gamma_D^*-\delta_D^*\gamma.
\]

A common misconception is that adding more proxies automatically strengthens identification. The invalid-proxy formulation shows that this is not generally true: with many proxies, exclusion violations can introduce bias unless uniqueness restrictions such as the majority rule are satisfied [2507.19623].

## 3. Violations of completeness and \(U\)-relevance

A distinctive contribution of the proximal LSEM literature is the derivation of closed-form bias formulas when proximal identification assumptions fail [2208.00105]. Cobzaru et al. study a simplified setting with \(X\equiv 0\), \(\gamma_{au}=0\), \(p=2\), and \(m=n=1\), and analyze two minimal failures.

Under a completeness violation, let \(U=(U_1,U_2)^T\), let both \(Z\) and \(W\) depend on \(U_1\) and \(U_2\), and assume \(\theta_{u2}\mu_{u2}\neq 0\) so that \(\dim U>\dim Z,W\) and completeness fails. If one fits the wrong linear bridge
\[
h(w,a;b)=b_0+b_a a+b_w w+b_{aw}aw
\]
by Method-of-Moments, the asymptotic bias in the ATE estimate \(\hat\psi_{\rm POR}\) is
\[
\delta_{\rm POR}
=
\frac{E[AU_1]}{E[A](1-E[A])}
\frac{\theta_{u2}\mu_{u2}}{\theta_{u1}\mu_{u1}}
\left\{
\frac{(1-E[A])S_2}{\mu_{u1}+S_2\frac{\theta_{u2}}{\theta_{u1}}\mu_{u2}}
+
\frac{E[A]S_1}{\mu_{u1}+S_1\frac{\theta_{u2}}{\theta_{u1}}\mu_{u2}}
\right\}\gamma_{u1},
\]
where
\[
S_1=\frac{(1-E[A])^2}{(1-E[A])(1-E[AU_1^2])-E[AU_1]^2},
\qquad
S_2=\frac{E[A]^2}{E[A]E[AU_1^2]-E[AU_1]^2}.
\]

Under a \(U\)-relevance violation, completeness is retained but the proxies fail to “see” an outcome-relevant confounder \(U_2\). In that case,
\[
\delta_{\rm POR}
=
-
\frac{\bigl(E[AU_1^2](1-E[AU_1^2])-E[AU_1]^2\bigr)\,E[AU_2]}
{\bigl(E[A]E[AU_1^2]-E[AU_1]^2\bigr)\bigl((1-E[A])(1-E[AU_1^2])-E[AU_1]^2\bigr)}
\,\gamma_{u2}.
\]

These formulas make the source of bias explicit. \(\theta_{u2}\mu_{u2}\) measures the dependence of \(Z\) and \(W\) on the “extra” confounder \(U_2\); \(E[AU_1]\) and \(E[A]\) measure the confounding of treatment by \(U_1\); \(S_1\) and \(S_2\) are normalization constants reflecting how well \(U_1\) can be recovered from \((Z,W)\); and in the \(U\)-relevance failure, \(\gamma_{u2}\) and \(E[AU_2]\) enter linearly [2208.00105].

Because completeness and \(U\)-relevance are empirically untestable, these closed forms provide a direct basis for quantitative bias analysis. The proposed workflow is to estimate or fix observable terms, specify plausible ranges for latent structural parameters, compute the implied range of \(\delta_{\rm POR}\), and, if desired, place priors on the latent parameters and propagate uncertainty to a posterior on \(\delta_{\rm POR}\) [2208.00105].

## 4. Penalized estimation and adaptive proxy selection

When some treatment-inducing proxies may be invalid, a direct estimator targets \((\beta,\alpha,\gamma)\) through penalized GMM:
\[
\min_{\beta,\alpha,\gamma}
\left\|P_M\!\left[Y-D\beta-Z\alpha-W\gamma\right]\right\|_2^2
+\lambda\|\alpha\|_1,
\]
where \(P_M\) projects onto the column space of \(M=[Z\mid D]\) and only \(\alpha\) is \(\ell_1\)-penalized [2507.19623].

An equivalent two-step implementation computes \(\widehat W=P_MW\), forms orthogonalized regressors
\[
\widetilde D=P_{\widehat W^\perp}D,\qquad
\widetilde Z=P_{\widehat W^\perp}Z,
\]
and solves
\[
\hat\alpha=
\arg\min_\alpha
\frac12\left\|Y-P_{\widetilde D^\perp}(\widetilde Z\alpha)\right\|_2^2+\lambda\|\alpha\|_1.
\]
Then
\[
\hat\beta=
\frac{\widetilde D^T\left[Y-\widetilde Z\hat\alpha\right]}{\|\widetilde D\|_2^2}.
\]
Because the LASSO shrinks some \(\hat\alpha_j\) to zero, it selects a subset of \(Z\)-variables treated as valid proxies [2507.19623].

To recover \(\gamma\), the model uses the ratio identity
\[
\gamma=\operatorname{median}_j(\Gamma_j^*/\delta_j^*),
\]
since
\[
\Gamma_j^*/\delta_j^*=\gamma+\alpha_j/\delta_j^*.
\]
Under the majority rule, the median of these ratios consistently estimates \(\gamma\) [2507.19623].

Standard LASSO may fail to recover the invalid set \(A\) unless the irrepresentable condition
\[
\|C_{A^c,A}C_{A,A}^{-1}\operatorname{sign}(\alpha_A)\|_\infty<1
\]
holds, where \(C=\operatorname{plim}(1/n)\widetilde Z^T\widetilde Z\). Because the \(\widetilde Z\)’s can remain highly correlated, the adaptive LASSO is proposed as a remedy. It first constructs a \(\sqrt n\)-consistent initial estimator using the median-of-ratios method,
\[
\gamma^m=\operatorname{median}(\widehat\Gamma_j/\widehat\delta_j),
\qquad
\alpha^m=\widehat\Gamma_Z-\widehat\delta_Z\gamma^m,
\]
and then solves
\[
\hat\alpha_{\mathrm{ad}}
=
\arg\min_\alpha
\left\|Y-P_{\widetilde D^\perp}(\widetilde Z\alpha)\right\|_2^2
+
\lambda_n\sum_{j=1}^{p_z}\frac{|\alpha_j|}{|\alpha_j^m|}.
\]
After selecting
\[
\widehat A_{\mathrm{ad}}=\{j:\hat\alpha_{\mathrm{ad},j}\neq 0\},
\]
the method treats those variables as confounders and re-estimates \(\beta\) by two-stage least squares:
\[
\hat\beta_{\mathrm{post}}
=
\frac{D^TP_{[Z_{\widehat A_{\mathrm{ad}}}\mid W]^\perp}Y}
{D^TP_{[Z_{\widehat A_{\mathrm{ad}}}\mid W]^\perp}D}.
\]
This procedure is designed to jointly select valid proxies and estimate the causal effect [2507.19623].

## 5. Large-sample properties and many invalid outcome proxies

Under the majority rule \(s_z<p_z/2\), restricted isometry or related design conditions, and tuning satisfying \(\lambda_n\to\infty\) but \(\lambda_n/\sqrt n\to 0\), the adaptive procedure has oracle-type guarantees [2507.19623]. With probability tending to one, \(\widehat A_{\mathrm{ad}}=A\); \(\hat\alpha_{\mathrm{ad},A}\) is \(\sqrt n\)-consistent and asymptotically Normal; and
\[
\sqrt n(\hat\beta_{\mathrm{post}}-\beta)\to^d \mathcal N(0,\sigma_{\mathrm{or}}^2),
\]
the same limit as the oracle 2SLS estimator that knows \(A\) in advance. This supports Wald-style confidence intervals,
\[
CI_{1-\alpha}(\beta)
=
\left[
\hat\beta_{\mathrm{post}}
\pm
z_{1-\alpha/2}\hat\sigma_{\mathrm{or}}/\sqrt n
\right].
\]

The framework extends to settings with many candidate outcome-inducing proxies \(W_1,\dots,W_{p_w}\), some of which may also be invalid, provided that fewer than half are invalid [2507.19623]. The algorithm runs the entire estimation procedure once for each candidate \(W_j\), producing \(\hat\beta_{\mathrm{post}}^j\), and then aggregates by the median:
\[
\hat\beta^{(p_w)}=\operatorname{median}_j(\hat\beta_{\mathrm{post}}^j).
\]
Under the majority rule on the \(W\)’s, this estimator is \(\sqrt n\)-consistent. Its exact limit law is an order statistic of a Normal vector and is typically non-Gaussian unless \(p_w-s_w\) is odd, so inference is based on nonparametric subsampling with subsample size \(b\ll n\), for example \(b\approx n^{4/5}\) [2507.19623].

Simulations support these theoretical results, and the method is applied to assess the effect of right heart catheterization on 30-day survival in ICU patient [2507.19623]. A plausible implication is that proximal linear SEMs are especially useful in observational studies where multiple proxy candidates exist but exclusion validity is uncertain.

## 6. Relation to regularized path-analysis SEM and terminological scope

A distinct line of research studies linear SEM in the path-analysis sense,
\[
Y=AY+\varepsilon,\qquad (I-A)Y=\varepsilon,
\]
with covariance reproduction
\[
\Sigma=(I-A)^{-1}\Psi(I-A)^{-T},
\]
where \(A\) is a path-coefficient matrix with \(\operatorname{diag}A=0\) and \(\Psi\) is the residual covariance [1809.06156]. In that literature, prior zero constraints are encoded by an index set \(I_A\), the nonconvex equality
\[
\Sigma^{-1}=(I-A)^T\Psi^{-1}(I-A)
\]
is relaxed by introducing a block variable \(X\), and the resulting confirmatory and exploratory formulations become convex after replacing the quadratic equality by the LMI \(X\succeq 0\) together with \(0\preceq X_4\preceq \alpha I\) and \(P(X_2)=I\) [1809.06156].

The exploratory sparse-SEM problem adds an \(\ell_1\)-type penalty,
\[
-\log\det X_1+\operatorname{Tr}(SX_1)+2\gamma\|P^c(X_2)\|_1,
\]
and is solved by proximal splitting methods, specifically PPXA and ADMM [1809.06156]. Under a low-rank exactness condition, rank \(X=n\), the relaxed solution coincides with a solution of the original nonconvex equality problem; the paper also defines an \(\alpha\)-critical threshold \(\alpha_c=n/\operatorname{Tr}(S^{-1})\) and notes that numerically one picks \(\alpha\lesssim \lambda_{\min}(S)\) to encourage rank \(X=n\) [1809.06156]. PPXA and ADMM were tested up to \(n=500\), with PPXA converging in \(O(10^2)\) iterations and running in approximately \(5\) minutes at low accuracy for \(n=500\) on a standard desktop; the dominant cost is an \(O((2n)^3)\) eigendecomposition each iteration [1809.06156].

This optimization-based usage should not be conflated with proxy-based proximal causal inference. In [1809.06156], “proximal” refers to proximal operators and splitting algorithms; in [2208.00105] and [2507.19623], it refers to identification through negative-control or proxy variables. The distinction is consequential. The proxy-based literature is concerned with hidden confounding, exclusion restrictions, completeness, and \(U\)-relevance, whereas the convex SEM literature is concerned with sparsity, low-rank exactness, and scalable estimation of path structures. Both operate within linear SEMs, but they address different inferential problems.

A further misconception is that proximal methods eliminate unverifiable assumptions. The bias-analysis results show the opposite: proximal inference still relies on empirically untestable conditions, and when completeness or \(U\)-relevance fails, the induced bias can be characterized but not ignored [2208.00105]. Conversely, the sparse path-analysis literature shows that convexification and proximal optimization can stabilize estimation and improve scalability, but they do not replace the causal assumptions required for proxy-based identification [1809.06156].

Source: https://www.emergentmind.com/topics/proximal-linear-structural-equations-model