Single Exposure Mediation Analysis Overview
- SE-MA is a method for decomposing the total causal effect of a single exposure into natural direct and indirect effects via mediators.
- It employs canonical data structures (Y, E, M, X) and counterfactual estimands to clearly define and quantify mediation pathways.
- Estimation strategies include regression models, g-computation, and semiparametric techniques, ensuring robust inference under key identification assumptions.
Single Exposure Mediation Analysis (SE-MA) denotes mediation analysis with one exposure of interest and an outcome linked by one or more mediators. In one usage, especially in exposure-mixture research, SE-MA is the baseline strategy that analyzes each exposure one at a time with standard mediator and outcome models; in the broader causal-inference literature, the same setting is formulated as a single point or baseline exposure or , a mediator or mediator vector , an outcome , and pre-exposure covariates or (Wang et al., 13 Sep 2025, Tchetgen et al., 2012, Steen et al., 2018). The central objective is to decompose the total causal effect of that single exposure into direct pathways not operating through the mediator system and indirect pathways that do.
1. Scope and canonical data structures
A canonical SE-MA data structure is , where is a single binary exposure measured at one time, 0 is measured after 1 and before 2, and 3 contains pre-exposure covariates (Tchetgen et al., 2012). Other formulations replace 4 by 5, allow 6 to be multivariate, and permit continuous exposures, binary outcomes, or survival outcomes. One general framework orders the variables as
7
with 8 for covariates, 9 for exposures, 0 for mediators, and 1 for the response; specializing to SE-MA amounts to setting the exposure dimension to one (Long et al., 2020).
SE-MA is therefore not restricted to the simplest one-exposure, one-mediator, linear-Gaussian setting. The literature represented here includes: single mediator models with natural direct and indirect effects; models with multiple, possibly causally dependent mediators; semiparametric formulations for a single binary exposure; and high-dimensional mediator settings in which the exposure is still singular but the mediator vector is large (Steen et al., 2018, Sun et al., 2020, Song et al., 2020).
In exposure-mixture studies, SE-MA has a narrower operational meaning. The exposure vector is 2, but the analyst selects one component 3 and fits a standard mediation model for that component, with or without including the remaining exposures 4 as covariates (Wang et al., 13 Sep 2025). This usage preserves the single-exposure estimand for 5, but it does not by itself define a global mediated effect of the full mixture.
2. Counterfactual estimands and effect decompositions
The standard SE-MA estimands are defined with counterfactuals. For a binary exposure 6, let 7 denote the mediator under 8, and let 9 denote the outcome under joint intervention on exposure and mediator. The natural indirect effect and natural direct effect use the nested counterfactual 0 (Steen et al., 2018). On the mean scale,
1
2
and by composition 3,
4
An equivalent semiparametric presentation writes the key mediation functional as 5, with 6 and 7, where 8 (Tchetgen et al., 2012).
SE-MA also includes controlled direct effects. For a fixed mediator level 9,
0
which contrasts exposure levels while holding the mediator fixed for everyone (Steen et al., 2018). Controlled effects do not require the same cross-world counterfactual 1 that natural effects do, and they therefore play a distinct role in designs aimed at realistic interventions.
With multiple manipulable mediators, one clinically oriented extension defines, for mediator 2,
3
and from this constructs the controlled direct effect 4, the controlled indirect effect 5, and the scaled controlled indirect effect
6
For each mediator 7, the total effect satisfies the exact decomposition
8
and averaging over 9 yields a corresponding average decomposition across mediators (Sun et al., 2020).
More generally, SE-MA encompasses path-specific effects. For a set of directed paths 0 from 1 to 2, the relevant counterfactual is written 3, with 4 set to 5 along paths in 6 and to 7 along the complement (Steen et al., 2018). This allows direct, indirect, and finer path-restricted contrasts to be defined within a common graphical framework.
3. Identification assumptions and graphical foundations
SE-MA identification rests on consistency, positivity, and no unmeasured confounding assumptions stated relative to the chosen estimand. In the single-exposure mediation tutorial for mixtures, the standard conditions are consistency and composition, positivity for exposure and mediator, SUTVA, no unmeasured exposure–outcome confounding, no unmeasured exposure–mediator confounding, no unmeasured mediator–outcome confounding conditional on exposure and covariates, and the cross-world independence condition
8
for natural effects (Wang et al., 13 Sep 2025). The semiparametric theory for a single binary exposure adopts the corresponding sequential ignorability conditions
9
and
0
together with positivity and consistency (Tchetgen et al., 2012).
Graphical models sharpen these assumptions by encoding when path-specific effects are identified and when they are not. In the simplest single-mediator setting, if a baseline covariate set 1 blocks all backdoor paths between 2 and 3 and is not affected by 4, the mediation formula identifies 5 from observed data (Steen et al., 2018). However, the same source emphasizes that adjustment for confounding is insufficient for identification of path-specific effects because their magnitude also depends on cross-world dependencies between exposure effects on the mediator and mediator effects on the outcome.
Two nonidentification mechanisms recur throughout the literature. The first is unmeasured mediator–outcome confounding, often represented by a latent variable 6 affecting both 7 and 8; this blocks identification of 9 even when treatment is randomized (Steen et al., 2018). The second is an exposure-induced mediator–outcome confounder, typically a variable 0 on a path 1; adjusting for 2 blocks part of the mediated pathway, while failing to adjust leaves mediator–outcome confounding (Steen et al., 2018, Tchetgen et al., 2012).
For multi-mediator SE-MA, the graphical criteria become more stringent. The recanting witness and recanting district criteria characterize when path-specific effects are impossible to identify because a node or district would need incompatible exposure assignments along different outgoing paths (Steen et al., 2018). A related practical point arises in exposure-mixture applications: if co-exposures 3 affect both mediator and outcome and are associated with the focal exposure 4, omitting them makes SE-MA “without co-exposures” not causally interpretable (Wang et al., 13 Sep 2025).
4. Statistical models and estimation strategies
The classical regression formulation of SE-MA is the Baron–Kenny linear model. With one exposure 5, one mediator 6, and covariates 7,
8
9
while the short regressions that omit 0 deliver the usual direct-effect estimate 1 and indirect-effect estimate 2 (Zhang et al., 2022). In the no-unmeasured-confounding linear setting, the product and difference methods coincide.
The standard parametric SE-MA formulas used in the mixture tutorial are obtained from linear models without exposure–mediator interaction: 3
4
For contrast 5,
6
7
8
When co-exposures are adjusted, the same coefficient formulas apply, but now as effects of 9 conditional on 0 (Wang et al., 13 Sep 2025).
More general SE-MA estimators use g-computation or semiparametric influence-function methods. One multivariate-mediator framework defines
1
then obtains direct, indirect, and total effects from contrasts among 2, 3, and 4 on the mean, odds-ratio, or restricted-mean-survival-time scale (Long et al., 2020). Estimation proceeds by fitting a multivariate linear mediator model, an outcome model appropriate to continuous, binary, or survival 5, simulating mediators from the fitted mediator model, and averaging predicted outcomes over the empirical covariate distribution. In linear single-mediator models, this Monte Carlo procedure converges to the usual path-coefficient formulas (Long et al., 2020).
The semiparametric theory for a single binary exposure develops efficient influence functions for 6, NDE, and NIE, and proposes a triply robust estimator of 7 that is consistent and asymptotically normal under the union model in which any two of the three nuisance components—outcome regression, mediator density, and exposure propensity—are correctly specified (Tchetgen et al., 2012). At the intersection of those nuisance models, the resulting NDE and NIE estimators achieve the nonparametric efficiency bound (Tchetgen et al., 2012). This places SE-MA on the same semiparametric footing as modern total-effect estimation.
5. Multiple mediators, high-dimensional mediators, and mixture settings
Single exposure does not imply single mediator. One framework for multivariate mediators and nonlinear outcomes explicitly treats 8 as a joint mediator block and defines a joint NIE for the entire vector rather than path-specific effects for individual mediators, noting that with correlated mediators and exposure-induced mediator–outcome confounding, path-specific effects are generally not identified (Long et al., 2020). A plausible implication is that “single exposure” is best regarded as a restriction on the treatment dimension, not on mediator dimensionality.
High-dimensional SE-MA has motivated mediator-selection methods that target the natural indirect effect itself. In the Bayesian sparse framework, the mediator-specific indirect contribution is
9
where 00 is the exposure–mediator effect and 01 is the mediator–outcome effect (Song et al., 2020). Two priors are proposed: a four-component Gaussian mixture prior, which explicitly encodes the composite null states 02, 03, 04, and 05; and a product threshold Gaussian prior, which thresholds coefficients according to marginal and product magnitudes (Song et al., 2020). Posterior inclusion probabilities (PIPs) serve as mediator-selection statistics, and PIP 06 is used as the default cutoff (Song et al., 2020).
In exposure-mixture research, SE-MA is the baseline comparator for mixture-specific mediation methods such as principal component mediation analysis, environmental risk score mediation analysis, and Bayesian kernel machine regression causal mediation analysis (Wang et al., 13 Sep 2025). The tutorial is explicit that unadjusted SE-MA is structurally misspecified in mixtures: average relative bias exceeds 345% under weak mediation and is about 400% under strong mediation, and the false positive rate ranges from 0.34 to 0.59 (Wang et al., 13 Sep 2025). Co-exposure-adjusted SE-MA substantially reduces bias—below 25% across all simulated scenarios and below 10% when 07 and 08—but becomes conservative, with false positive rate below 0.02 and true positive rate as low as about 0.03 when 09 and 10 (Wang et al., 13 Sep 2025). The main limitation is that SE-MA remains exposure-specific and does not, in general, identify a coherent global indirect effect of the mixture (Wang et al., 13 Sep 2025).
6. Sensitivity analysis, robustness, and common limitations
A central limitation of natural-effect SE-MA is its dependence on untestable assumptions, especially mediator ignorability and cross-world independence. One semiparametric sensitivity framework parameterizes violations of mediator ignorability through the selection bias function
11
which is zero under mediator ignorability and can be varied over a sensitivity model 12 to study how NDE estimates change (Tchetgen et al., 2012). This yields a doubly robust sensitivity estimator of the NDE when the mediator density model is correct and the chosen 13 is treated as fixed (Tchetgen et al., 2012).
Within the Baron–Kenny linear setting, unmeasured confounding can also be parameterized through partial correlations aligned with the mediation DAG: 14 These measure the partial correlation between the unmeasured confounder and the exposure, mediator, and outcome, respectively (Zhang et al., 2022). The same work defines the robustness value for mediation as the minimum value of the maximum proportion of variability explained by the unmeasured confounding, for the exposure, mediator, and outcome, needed to overturn the direct- or indirect-effect conclusion (Zhang et al., 2022). The bounds are proved attainable and thus