Papers
Topics
Authors
Recent
Search
2000 character limit reached

Proximal Path-Specific Inference

Published 10 May 2026 in stat.ME, math.ST, and stat.ML | (2605.09462v1)

Abstract: Causal mediation analysis has been extended to estimate path-specific effects with multiple intermediate variables, isolating treatment effects through a mediator of interest while excluding pathways through its ancestors. Such analyses address bias from recanting witnesses, i.e., treatment-induced mediator-outcome confounders. However, existing methods typically rely on stringent assumptions precluding general unmeasured confounding, which are often violated in practice. In this paper, we relax these restrictions by leveraging observed covariates as proxy variables to accommodate unmeasured confounding among the treatment, recanting witness, mediator, and outcome. Using proximal confounding bridge functions, we develop four nonparametric identification strategies for the path-specific effect. We further derive the efficient influence function and propose a quadruply robust, locally efficient estimator. To handle high-dimensional nuisance parameters, we propose a proximal debiased machine learning approach. We theoretically guarantee that our estimator achieves n\sqrt{n}-consistency and asymptotic normality even when machine learning estimators for nuisance functions converge at slower rates. Our approaches are validated via semiparametric and nonparametric simulations and an application to the CDC WONDER Natality study, estimating the path-specific effect of prenatal care on preterm birth through preeclampsia, independent of maternal smoking during pregnancy.

Authors (4)

Summary

  • The paper proposes a novel approach to isolate indirect effects and address bias from recanting witnesses using proxy variables.
  • It employs nested Fredholm integral equations to construct confounding bridge functions under minimal identification assumptions.
  • The methodology is validated through quadruply robust, semiparametric, and debiased machine learning estimators ensuring local efficiency.

Proximal Path-Specific Inference: A Formal and Technical Summary

Motivation and Problem Formulation

The paper "Proximal Path-Specific Inference" (2605.09462) addresses the identification and estimation of path-specific causal effects in mediation analysis under scenarios where both recanting witnesses (treatment-induced mediator-outcome confounders) and general unmeasured confounding are present. Traditional approaches to mediation analysis, including natural indirect effects and interventional indirect effects, are heavily reliant on restrictive assumptions such as sequential ignorability or no unmeasured mediator-outcome confounding. These assumptions often fail in epidemiological and social science contexts, particularly when mediators and recanting witnesses are treatment-induced and affected by latent confounders unaccounted for in the observed data.

The primary estimand considered is the mean difference in potential outcomes under an intervention where the mediator MM is set according to its counterfactual distribution under A=0A=0 but the recanting witness DD follows its natural value under A=1A=1: PAMY=E[Y(1,D(1),M(1,D(1)))]−E[Y(1,D(1),M(0,D(1)))]\mathcal{P}_{AMY} = E\left [ Y\left ( 1, D(1), M\left ( 1, D(1) \right ) \right ) \right ] - E\left [ Y\left ( 1, D(1), M\left ( 0, D(1) \right ) \right ) \right ] This estimand isolates the indirect effect through MM, holding DD fixed, and excludes pathways via recanting witnesses, thus resolving the classical bias from treatment-induced confounding.

The challenge is compounded by latent confounders UU affecting AA, DD, A=0A=00, and A=0A=01 (Figure 1). These confounders invalidate standard identification strategies, necessitating novel methodology capable of robust inference under minimal assumptions.

(Figure 1)

Figure 1: Causal graphs contrasting restricted (A=0A=02 only affecting A=0A=03 and A=0A=04) and generalized (A=0A=05 affecting A=0A=06, A=0A=07, A=0A=08, and A=0A=09) unmeasured confounding structures; the latter captures realistic epidemiologic data-generating mechanisms.

Identification via Proximal Proxy Variables

To address unmeasured confounding, the authors leverage the proximal causal inference framework, utilizing observed covariates as proxy variables for latent confounders. Proxy variables DD0 (treatment-inducing) and DD1 (outcome-inducing) are selected based on structural criteria: DD2 is a cause of DD3 associated with DD4, DD5, DD6 only via DD7, DD8, and DD9; A=1A=10 is a cause of A=1A=11 associated with A=1A=12, A=1A=13, A=1A=14 solely via A=1A=15, A=1A=16. This structure is formalized in Assumption 5 and illustrated graphically (Figure 2).

The core identification strategy involves constructing and solving a sequence of nested Fredholm integral equations to obtain confounding bridge functions. Four complementary strategies are developed, relying on outcome or treatment bridge functions or their hybrids. Completeness conditions for proxy variables are required, ensuring that variation in proxies adequately captures the variability in latent confounders. Figure 2

Figure 2: Comparison of bias across sample sizes (n=200, n=1000, n=2000) for estimation methods (por, pipw, phybrid_1, phybrid_2, and P-DML). The dashed red line indicates zero bias.

Semiparametric Estimation and Quadruply Robust Methods

Building upon the identification results, efficient influence functions are derived for the path-specific estimand. The authors develop semiparametric estimators based on these influence functions, yielding the following main features:

  • Quadruply Robustness: The estimator remains consistent if any one out of four specified unions of model sets for bridge functions is correctly specified.
  • Local Efficiency: When all nuisance models are correctly specified, the estimator achieves the lowest semiparametric variance.
  • Neyman Orthogonality: The estimator is insensitive to slow convergence of nuisance function estimators, retaining valid inference under machine learning-based bridge estimation.

The construction of estimators utilizes a sequence of estimating equations rather than explicit regression modeling, thus avoiding standard pitfalls associated with plug-in ML-based mediation estimators and enabling nonparametric learning of high-dimensional nuisance structures.

Debiased Machine Learning Approach

For high-dimensional nuisance parameter settings, the paper introduces a proximal debiased machine learning (P-DML) procedure using cross-fitting and regularized minimax learning in RKHS. The estimator is constructed from empirical averages of the efficient influence function, with foldwise nuisance function estimation employing minimax estimation over conditional moment equations. This procedure achieves A=1A=17-consistency and asymptotic normality under mild regularity conditions, even when nuisance estimates converge at A=1A=18 rates, aligning with contemporary double machine learning theory.

The estimator exhibits low bias and mean squared error across a wide range of simulation scenarios, maintaining accuracy even in settings with severe model misspecification (Figure 2, Table 1). The theoretical guarantees are established via second-order expansions and variance bounds for the efficient influence function.

Empirical Evaluation and Application

Two sets of numerical experiments are presented:

  1. Semiparametric Simulations: Multiple estimators (P-OR, P-IPW, P-hybridA=1A=19, P-hybridPAMY=E[Y(1,D(1),M(1,D(1)))]−E[Y(1,D(1),M(0,D(1)))]\mathcal{P}_{AMY} = E\left [ Y\left ( 1, D(1), M\left ( 1, D(1) \right ) \right ) \right ] - E\left [ Y\left ( 1, D(1), M\left ( 0, D(1) \right ) \right ) \right ]0, P-quadR) are compared under varying model misspecification, confirming robustness claims. P-quadR consistently achieves low bias and nominal coverage when only one model set is correctly specified.
  2. Nonparametric Simulations: The P-DML estimator is benchmarked against other proximal strategies using RKHS-based regularized minimax learning. P-DML demonstrates superior statistical efficiency and bias control; bias converges quickly with increasing sample size.

The methods are further validated on CDC WONDER Natality data, where path-specific effects of prenatal care on preterm birth mediated by preeclampsia (excluding smoking pathways) are estimated. Proxy variables (marital status and pre-pregnancy hypertension) are selected based on domain considerations. The quadruply robust estimator identifies a small but statistically significant reduction in preterm birth risk specific to the PAMY=E[Y(1,D(1),M(1,D(1)))]−E[Y(1,D(1),M(0,D(1)))]\mathcal{P}_{AMY} = E\left [ Y\left ( 1, D(1), M\left ( 1, D(1) \right ) \right ) \right ] - E\left [ Y\left ( 1, D(1), M\left ( 0, D(1) \right ) \right ) \right ]1 pathway, with confidence intervals narrower than those produced by alternative estimators and values consistent with interventional estimates reported previously.

Implications and Future Directions

This research generalizes proximal causal inference approaches to path-specific mediation analysis in the presence of recanting witnesses and pervasive unmeasured confounding. The quadruply robust and debiased machine learning methods substantially relax the required identification assumptions, offering practitioners a principled framework to estimate isolable indirect causal pathways in realistic observational studies.

Theoretically, the framework connects causal mediation, instrumental variable identification, and confounding bridge estimation, and sets the foundation for future work in:

  • High-dimensional variable selection for proxy construction
  • Extensions to longitudinal and dynamic mediation structures
  • Machine learning-based bridge estimation in broader causal architectures
  • Applications in genomic, social, and policy research where unmeasured confounding is unavoidable

Conclusion

"Proximal Path-Specific Inference" provides a rigorous methodological foundation for estimating path-specific effects with complex unmeasured confounding and recanting witnesses, leveraging proxy variables and efficient influence function-based estimation. The technical contributions encompass nonparametric identification, quadruply robust estimation, and debiased machine learning, validated by simulations and application to large epidemiological datasets. The practical and theoretical implications are substantial, positioning this framework as a preferred solution for mediation analysis under realistic assumptions in contemporary causal inference.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.