---
title: Proximal Causal Inference (PCI)
url: https://www.emergentmind.com/topics/proximal-causal-inference-pci
type: topic
---

# Proximal Causal Inference (PCI)

Searching arXiv for recent and foundational papers on proximal causal inference to ground the encyclopedia entry.
Proximal causal inference (PCI) is a causal inference framework for identifying and estimating causal effects in the presence of unmeasured confounding by using observed proxy variables rather than requiring that all confounders be directly measured. In the canonical point-treatment formulation, PCI introduces an unobserved confounder \(U\), observed baseline covariates \(X\), treatment-inducing confounding proxies \(Z\), and outcome-inducing confounding proxies \(W\), and replaces standard exchangeability \(Y(a)\perp A\mid X\) with latent unconfoundedness \(Y(a)\perp A\mid U,X\) together with proxy restrictions such as \(Z \perp Y \mid U,A,X\) and \(W \perp (A,Z)\mid U,X\) [2011.08411]. Identification proceeds through confounding bridge functions that solve inverse conditional moment equations, yielding the proximal g-formula and related estimating equations for average treatment effects and other causal functionals [2009.10982; 2011.08411]. Subsequent work has extended this core idea to longitudinal treatment regimes, synthetic controls, hidden mediators, modified treatment policies, survival outcomes, invalid proxies, and non-unique bridge settings [2109.07030; 2108.13935; 2111.02927; 2512.12038; 2409.08924; 2506.13152; 2303.10134].

## 1. Conceptual formulation and proxy structure

PCI addresses the observational setting in which treatment \(A\), outcome \(Y\), and measured covariates are observed, but the standard “no unobserved confounding” assumption is not credible because the true confounding mechanism is only partially observed. A common formulation partitions measured covariates as \(L=(X,Z,W)\), where \(X\) contains observed common causes, \(Z\) contains treatment-inducing confounding proxies, and \(W\) contains outcome-inducing confounding proxies, while \(U\) denotes an unobserved confounder such that \(Y(a)\perp A\mid U,X\) [2011.08411; 2009.10982; 2512.24413].

The defining proxy restrictions are asymmetric. In the point-treatment proximal setup, treatment confounding proxies satisfy
\[
Z \perp Y \mid U,A,X,
\]
while outcome confounding proxies satisfy
\[
W \perp (A,Z)\mid U,X.
\]
These conditions encode that \(Z\) is outcome-disconnected given \(U,A,X\), whereas \(W\) is treatment-disconnected given \(U,X\) [2011.08411; 2512.24413]. This negative-control-style asymmetry is central: PCI does not treat proxies as direct substitutes for \(U\), but instead uses their distinct conditional independence roles to recover causal functionals through bridge equations [2009.10982; 2401.06687].

The framework also requires consistency and positivity. In the point-treatment setting these are stated as
\[
Y=Y(A)\quad \text{a.s.},
\]
and
\[
0<\Pr(A=a\mid U,X)<1 \quad \text{a.s.},
\]
for \(a\in\{0,1\}\) [2011.08411]. This latent positivity condition is stronger than ordinary positivity given observed covariates, because support is required conditional on the unobserved confounder.

A recurrent practical interpretation is that measured variables are not the true confounders but noisy, incomplete, or indirect measurements of deeper latent processes such as disease severity, physiological status, treatment propensity, or health-seeking behavior [2009.10982; 2512.24413]. This suggests PCI is best understood not as a relaxation of assumptions in a purely formal sense, but as a replacement of standard exchangeability by a proxy-based latent structure.

## 2. Bridge functions and nonparametric identification

The foundational identification strategy uses an outcome confounding bridge \(h(W,A,X)\) satisfying
\[
E[Y\mid Z,A,X]=E[h(W,A,X)\mid Z,A,X].
\]
Under the proximal assumptions and completeness of \(Z\) for \(U\), this implies
\[
E[Y(a)] = \int_{\mathcal X}\int h(w,a,x)\, dF(w\mid x)\, dF(x),
\]
and therefore the average treatment effect is identified as
\[
\psi = E\{Y(1)-Y(0)\}=E[h(W,1,X)-h(W,0,X)].
\]
This is the proximal g-formula [2011.08411; 2009.10982].

A complementary identification route uses a treatment confounding bridge \(q(Z,a,X)\) satisfying
\[
E\{q(Z,a,X)\mid W,A=a,X\}=\frac{1}{f(A=a\mid W,X)}.
\]
Under the corresponding completeness condition for \(W\), this yields
\[
E\{Y(a)\}=E\bigl\{I(A=a)q(Z,a,X)Y\bigr\},
\]
and thus
\[
\psi = E\bigl\{AYq(Z,1,X)-(1-A)Yq(Z,0,X)\bigr\}.
\]
This is the proximal analog of inverse probability weighting [2011.08411].

The two bridge routes are structurally parallel: \(h\) plays the role of a latent outcome regression surrogate, while \(q\) plays the role of a latent inverse propensity surrogate. In both cases the underlying equations are Fredholm integral equations of the first kind, and identification depends on existence of bridge solutions and completeness conditions such as
\[
E\{g(U)\mid Z,A=a,X=x\}=0 \iff g(U)=0
\]
or
\[
E\{g(U)\mid W,A=a,X=x\}=0 \iff g(U)=0
\]
[2011.08411; 2303.10134].

A further conceptual development concerns non-uniqueness. PCI identification of a causal estimand can survive even when the bridge function itself is not uniquely identified. In operator notation, if \(T_o h = E[Y\mid Z,A,X]\) defines the outcome bridge equation, then any two solutions differ by an element of the null space \(\mathcal N(T_o)\); nevertheless a linear functional such as the counterfactual mean remains identified if its representer lies in \(\mathcal N(T_o)^\perp\) [2303.10134]. This result separates identification of the bridge from identification of the causal estimand.

## 3. Completeness, inverse problems, and semiparametric inference

Completeness is the technical condition that ensures the proxies are sufficiently informative about the hidden confounder or each other. In the review literature it is described as a richness or injectivity condition on conditional expectation operators, and in discrete settings it is often linked to rank conditions [2512.24413]. In linear specializations, completeness can reduce to full-rank requirements. For example, in the proximal synthetic control formulation, unique identification of synthetic control weights \(\bw\) follows from the full row rank of \(E[g(Z_t)\ww]\), which is described as the linear identification analogue of the rank/completeness conditions in standard PCI [2108.13935].

The semiparametric theory of PCI was developed for the average treatment effect in the point-treatment setting. Under a semiparametric model in which bridge functions exist and suitable regularity conditions hold, the efficient influence function for the ATE is
\[
EIF(\psi) = (-1)^{1-A} q(Z,A,X)\bigl[Y-h(W,A,X)\bigr] + h(W,1,X)-h(W,0,X)-\psi.
\]
This yields the semiparametric efficiency bound \(E\{EIF^2(\psi)\}\) [2011.08411]. The same work characterizes proximal outcome-regression, proximal inverse-probability-weighted, and proximal doubly robust estimators, with the doubly robust estimator
\[
\widehat \psi_{PDR} = \mathbb P_n\Bigl[ (-1)^{1-A}\widehat q(Z,A,X)\{Y-\widehat h(W,A,X)\} + \widehat h(W,1,X)-\widehat h(W,0,X) \Bigr].
\]
This estimator is consistent if either the outcome bridge model or the treatment bridge model is correctly specified, and is locally semiparametrically efficient when both are correctly specified [2011.08411].

The non-uniqueness literature extends this semiparametric perspective. When bridge solutions are set-valued rather than point-identified, one can estimate the entire solution set
\[
\mathcal H_0=\{h\in\mathcal H:\E[h(W,A,X)\mid Z,A,X]=\E[Y\mid Z,A,X]\},
\]
select a uniquely defined representative by minimizing a convex criterion, and then debias the plug-in estimator using a representer \(g_0\). The resulting estimator
\[
\widehat\mu_{a\text{-}db}=\widehat\mu_a+\widehat r_n(\widehat h_0)
\]
is root-\(n\) consistent and asymptotically normal under stated regularity and rate conditions [2303.10134].

A plausible implication is that PCI has evolved from a purely identification-oriented framework into a full inferential program with efficiency theory, doubly robust estimation, and orthogonalized or debiased estimators, while still retaining difficult untestable assumptions about proxies and completeness.

## 4. Longitudinal, panel, and structured-data extensions

PCI has been extended beyond point treatment to complex longitudinal treatment regimes. In the two-timepoint formulation of proximal causal inference for complex longitudinal studies, observed covariates are partitioned into \(\bar X(1)\), \(\bar Z(1)\), and \(\bar W(1)\), with latent time-varying confounders \(\bar U(1)\) and treatment history \(\bar A(1)\) [2109.07030]. Under sequential proxy restrictions and completeness assumptions, longitudinal outcome bridge functions \(H_1\{\bar a(1)\}\) and \(H_0\{\bar a(1)\}\) identify
\[
E\{Y_{\bar a(1)}\mid V\} = E\!\left[h_0\{W(0),\bar a(1),X(0)\}\mid V\right],
\]
while treatment bridge functions \(Q_0\{a(0)\}\) and \(Q_1\{\bar a(1)\}\) identify
\[
E\{Y_{\bar a(1)}\mid V\} = E\!\left[ Y\mathbbm{1}\{\bar A(1)=\bar a(1)\}Q_1\{\bar a(1)\} \mid V \right].
\]
This extension yields proximal analogs of longitudinal outcome regression, inverse weighting, and doubly robust estimation for marginal structural mean models [2109.07030].

In panel and comparative case-study settings, PCI has been used to reformulate synthetic control. In the synthetic control setting with one treated unit and donor outcomes \(\ww\), the paper “Theory for identification and Inference with Synthetic Controls: A Proximal Causal Inference Framework” interprets donor outcomes as proxies for latent common factors and unused controls as auxiliary proxies \(Z_t\) [2108.13935]. Under the proxy assumption
\[
Z_t \perp \{Y_t,\ww\}\mid \lambda_t,
\]
the synthetic control weights satisfy the proximal bridge restriction
\[
E\Big[Y_t-\sum_{i\in D}\alpha_i W_{it}\,\Big|\, Z_t\Big]=0,\qquad t\le T_0.
\]
This supports GMM identification of the synthetic weights and model-based inference for post-treatment effects, rather than relying solely on pre-treatment balancing or placebo permutations [2108.13935].

PCI has also been adapted to unstructured data. In text-based PCI, two instances of pre-treatment text are split into separate channels and mapped to proxies \(Z\) and \(W\) using zero-shot models:
\[
Z \gets {\cal M}_1({\bf T}^{\text{pre}_1}), \qquad W \gets {\cal M}_2({\bf T}^{\text{pre}_2}).
\]
Under the design condition
\[
{\bf T}^{\text{pre}_1} \perp {\bf T}^{\text{pre}_2 \mid U, {\bf C},
\]
the resulting proxies satisfy the canonical proximal assumptions, whereas using the same text passage for both proxies or using post-treatment text does not [2401.06687]. This line of work shows that PCI can be driven by design of proxy channels rather than only by structured-variable selection.

Another structured-data extension targets modified treatment policies for continuous exposures. In that setting, the target is
\[
\psi_0=\mathbb{E}\!\left[Y\{q(A,L)\}\right],
\]
and treatment-bridge identification involves a policy-specific Jacobian-density correction
\[
\alpha_0(a,l,w) :=I_q(a,l)\frac{dq^{-1}(a,l)}{da} \frac{p_{A|L,W}\{q^{-1}(a,l)\mid l,w\}}{p_{A|L,W}(a\mid l,w)}.
\]
Under proximal assumptions and either an observed outcome bridge or an observed treatment bridge, the mean under the modified treatment policy is identified by
\[
\psi_0 = \mathbb E\left[h_0\{q(A,L),L,W\}\right] = \mathbb E\left\{Yg_0(A,L,Z)\right\}
\]
[2512.12038]. This is a distinct extension because the bridge now must encode the policy map \(q\), not merely static interventions.

## 5. Mediation, hidden outcomes, survival, and decision learning

PCI has been generalized from hidden confounding to hidden mediators. In “Causal Inference with Hidden Mediators,” the mediator \(M\) is unobserved but proxies \(Z\) and \(W\) are observed, with proxy restrictions
\[
Y \perp Z \mid \{A,M,X\}, \qquad W \perp \{A,Z\} \mid \{M,X\}.
\]
Under mediator-specific bridge and completeness assumptions, the hidden mediation functional
\[
\theta_{MED}^{a',a} = E\!\left[Y^{(a',M^{(a)})}\right]
\]
is identified through outcome-bridge and treatment-bridge analogues [2111.02927]. The same framework yields a hidden front-door criterion and identification of population intervention indirect effects [2111.02927].

More recent mediation work uses PCI for path-specific effects in the presence of hidden recanting witnesses. In the sequential mediation setup with hidden \(M_1\), observed downstream mediator \(M_2\), and proxies \(Z\) and \(W\), the target
\[
\psi = E\!\left[Y\big(M_2(M_1(0),1),\,M_1(0),\,0\big)\right]
\]
is identified through proximal outcome-regression, hybrid, and inverse-weighted representations [2606.17600]. The efficient influence function and a proximal multiply robust estimator are derived, with consistency if at least one of several nuisance-model combinations is correctly specified [2606.17600].

PCI has also been extended to hidden outcomes rather than hidden confounders. In that setting, the outcome \(Y\) is never observed; instead three proxies \(W,Z,V\) for \(Y\) are observed. Under conditional mutual independence of the proxies given \(Y,A,C\), completeness, and label-identification assumptions, the full-data law \(p(A,Y,C,W,Z,V)\) is identified [2605.09849]. The paper then derives an observed-data influence function for \(\psi_a=\mathbb E[Y(a)]\),
\[
\phi_{\mathrm{obs}}(O;\psi_a,\eta) = \frac{\mathbb{I}(A=a)}{f(A\mid C)} \left\{ \sum_{y\in\mathcal Y_{A,C}} \omega_y(O;\eta)\, y - \mathbb{E}[Y\mid A=a,C] \right\} + \mathbb{E}[Y\mid A=a,C]-\psi_a,
\]
and establishes multiple robustness under six nuisance-model configurations [2605.09849]. This is adjacent to PCI rather than canonical double-negative-control PCI, but it extends proxy-based causal identification into a latent-outcome regime.

Survival analysis has likewise been incorporated. In regression-based PCI for right-censored time-to-event data, the primary survival endpoint follows an additive hazards model
\[
\lambda_T(t\mid A,U,X)=\beta_0(t)+\beta_A^T A+\beta_X^T X+\beta_U U.
\]
Assuming a location-shift model for \(U\),
\[
U=E(U\mid A,Z,X)+\epsilon, \qquad E(U\mid A,Z,X)=\gamma_0+\gamma_A^T A+\gamma_Z^T Z+\gamma_X^T X,
\]
the fitted first-stage regression for a negative control outcome \(W\) serves as a surrogate confounding score, yielding a two-stage regression procedure called Proximal Two-Stage Least-Squares for Survival data (P2SLS-Surv) [2409.08924]. The method is developed for continuous, count, and right-censored time-to-event negative control outcomes.

PCI has further entered decision learning. In “Optimal Treatment Regimes for Proximal Causal Learning,” the observed covariates decompose as \(L=(X,W,Z)\), and PCI bridge functions are used to identify treatment regime value functions despite latent confounding [2212.09494]. The paper defines a richer regime class
\[
\mathcal D_{\mathcal{ZW}^{\Pi}}
\]
that adaptively switches between a \(Z\)-based proximal regime and a \(W\)-based proximal regime using an individualized selector \(\pi(X)\), and proves
\[
V(d_{zw}^{\bar\pi})\ge \max\{V(d_z),V(d_w)\}.
\]
This shows that PCI supports not only effect estimation but also policy optimization under hidden confounding [2212.09494].

## 6. Robustness to invalid proxies, applications, and ongoing challenges

A major practical challenge is that validity of proxies is typically untestable. Standard PCI assumes that the analyst-specified \(Z\) and \(W\) sets satisfy their exclusion restrictions, but recent work relaxes this assumption. In “Fortified Proximal Causal Inference with Many Invalid Proxies,” the analyst observes \(K\) candidate treatment proxies \(Z_1,\dots,Z_K\), of which at least \(\gamma\) are valid, without knowing which ones [2506.13152]. Identification is obtained through the function space
\[
H_\gamma := \left\{d\in L_2(Z,A,X): E\{d(Z,A,X)\mid Z_{-\mathcal J},X\}=0,\ \forall \mathcal{J}\in P_{\gamma}([K])\right\},
\]
and fortified bridge moments such as
\[
E[d(Z,A,X)\{Y-h^*(W,A,X)\}]=0,\qquad \forall d\in H_\gamma.
\]
This yields an efficient influence function and a fortified proximal multiply robust estimator under the union model \(M_{1,\gamma}\cup M_{2,\gamma}\cup M_{3,\gamma}\) [2506.13152].

A related linear-proximal development studies adaptive PCI with some invalid proxies. In a canonical linear model
\[
\mathbb{E}(Y \mid D, \mathbf Z, U) = \beta D + \boldsymbol{\alpha}^{\intercal}\mathbf Z + \beta_u U, \qquad \mathbb{E}(W \mid D, \mathbf Z, U) = \eta_u U,
\]
a candidate treatment proxy is valid iff its direct-effect coefficient \(\alpha_j\) is zero [2507.19623]. Under majority validity of the candidate treatment proxies, a median estimator for the bridge coefficient \(\gamma=\beta_u/\eta_u\) and an adaptive LASSO on \(\boldsymbol\alpha\) select invalid proxies and recover an oracle-equivalent post-selection estimator \(\widehat\beta_{post}\) [2507.19623]. This suggests that robustness to proxy invalidity is becoming a central theme in PCI methodology.

Applied work illustrates both the promise and fragility of the framework. The foundational semiparametric PCI paper reanalyzed the SUPPORT study on right heart catheterization and found proximal estimates of the 30-day survival effect more harmful than standard doubly robust adjustment, using
\[
Z=(\text{pafi1},\text{paco21}), \qquad W=(\text{ph1},\text{hema1})
\]
as proxies for latent physiological severity [2011.08411]. The survival-specific PCI paper also studied right heart catheterization, using \(\text{PaO}_2/\text{FiO}_2\) and \(\text{PaCO}_2\) as negative control exposures and blood pH and hematocrit as negative control outcomes, with a proximal estimate
\[
\hat\beta_A = 0.161 \; (95\%\,\text{CI } 0.025,\; 0.297)
\]
for increased mortality under the additive hazards model [2409.08924]. In synthetic control, PCI was applied to West Germany’s post-reunification GDP, yielding a proximal constant treatment effect estimate of \(-1200\) USD with heteroskedasticity-consistent and HAC confidence intervals [2108.13935]. In vaccine immunobridging, proximal modified-treatment-policy analysis found that upward shifts in Day 29 neutralizing antibody titer reduced estimated COVID risk and increased vaccine efficacy [2512.12038]. In fairness-oriented mediation analysis, PCI was used to estimate controlled direct effects of sociodemographic attributes on diagnosis decisions in UK Biobank, interpreting the direct effect as a bias-related pathway under latent health mediation assumptions [2501.16399].

The recurring limitations are also clear across the literature. Proxy validity, completeness, bridge existence, and latent positivity remain strong and largely untestable assumptions [2512.24413; 2011.08411]. Bridge estimation is often an ill-posed inverse problem, so practical methods rely on parametric models, RKHS regularization, adversarial learning, or sieve approximations [2011.08411; 2512.12038; 2303.10134]. Weak proxies can induce instability analogous to weak instruments [2409.08924]. Several papers explicitly note that choosing proxies requires substantive articulation of the hidden confounding mechanism rather than purely empirical screening [2512.24413].

These developments suggest that PCI is best understood as a family of proxy-based causal identification and inference methods anchored by bridge equations, rather than a single estimator or modeling recipe. Its core contribution is to formalize causal learning when measured covariates are acknowledged to be imperfect proxies of latent causal structure, and its recent literature shows a broadening from average treatment effects to survival, mediation, policy learning, panel data, text, hidden outcomes, and settings with proxy invalidity [2009.10982; 2011.08411; 2109.07030; 2111.02927; 2401.06687; 2409.08924; 2506.13152; 2512.12038].

Source: https://www.emergentmind.com/topics/proximal-causal-inference-pci