Papers
Topics
Authors
Recent
Search
2000 character limit reached

Position: A Potential Outcomes Perspective on Pearl's Causal Hierarchy

Published 28 Jan 2026 in stat.OT | (2601.20405v1)

Abstract: Pearl's causal hierarchy has garnered sustained attention as a foundational lens for formulating and understanding causal questions, and has been extensively discussed within the framework of structural causal models. In this paper, we revisit the hierarchy from a potential outcomes perspective and provide a formal, systematic classification of how various causal estimands are mapped to specific layers. Building on this classification, we summarize key identifiability challenges for estimands at different layers and review general strategies for achieving identification under varying assumptions. Our perspective is both intuitive and theoretically grounded, as higher layers of the hierarchy correspond to progressively richer features of the potential outcomes distribution, which in turn require stronger assumptions for identification. We expect this perspective to help clarify and deepen understanding of various causal estimands, particularly those in the third layer of the causal hierarchy, along with their associated identifiability challenges, identifiability strategies, and application scenarios.

Authors (2)

Summary

  • The paper offers a categorization of causal estimands using the potential outcomes framework, determining their layer on Pearl's hierarchy (dependency only on marginal or joint distributions or both).
  • It identifies intervention-based causal estimands (e.g., ATT, QTE, DRF) that require assuming marginal distribution of outcomes, while cross-world and individual-level estimands demand knowledge of joint distributions
  • The second layer estimands pose challenge with confounding between treatment and outcome and recommend consideration of auxiliary variables, data fusion, or sensitivity analysis for their identification.

Motivation and central claim

Pearl's causal hierarchy—association, intervention, counterfactuals—has been analyzed extensively within the structural causal model (SCM) framework, including from logical-probabilistic, inferential-graphical, and computational-complexity viewpoints. The position paper by Wu and Wang argues that a formal examination of the hierarchy through the lens of the potential outcomes framework of Neyman and Rubin has been missing, and proposes to fill this gap (2601.20405). The guiding question is practical: given a causal estimand, how can one determine which layer of the hierarchy it belongs to, and how can diverse causal estimands be systematically classified?

The paper's answer is a distributional criterion. An estimand belongs to the second layer (intervention) if it depends only on the marginal distributions of potential outcomes (conditional on observed treatment and covariates). An estimand belongs to the third layer (counterfactuals) if it additionally depends on the joint distribution of potential outcomes, involves nested potential outcomes under mutually exclusive interventions (cross-world relationships), or targets potential outcomes at the individual level. The first layer is omitted, on the grounds that causal inference concerns the second and third layers. The perspective is claimed to be both intuitive and theoretically grounded: higher layers correspond to progressively richer features of the potential outcomes distribution and therefore require stronger identification assumptions. This monotonicity is explicit—knowledge of individual-level counterfactuals implies knowledge of the joint distribution, which implies knowledge of the marginals, but not conversely. The framework thereby connects the SCM and potential outcomes traditions through the objects each must identify.

The second layer and its identification challenges

Second-layer estimands are functionals of marginal potential outcome distributions under a single intervention. Representative examples include the average treatment effect on the treated (ATT), the quantile treatment effect (QTE), the distributional treatment effect (DTE), the causal risk and odds ratios for binary outcomes, the dose–response function DRF(a)=E[Y(a)]\text{DRF}(a) = \mathbb{E}[Y(a)] for continuous treatments, and the average derivative effect. The paper also classifies counterfactual parity and the total effect in mediation analysis as second-layer estimands, a point that carries weight for practice: counterfactual parity, despite its "counterfactual" label, requires only marginal distributions.

The primary identification obstacle at this layer is confounding between treatment and outcome. The standard route is ignorability together with overlap; randomization suffices. When unmeasured confounding is a concern, the paper reviews auxiliary variables (instrumental variables, negative controls), data fusion across complementary sources, structural restrictions for multi-dimensional treatments or outcomes, and sensitivity analysis as a robustness tool.

Cross-world queries: the first sublayer of the third layer

The paper divides the third layer into two sublayers by inferential complexity. The first comprises cross-world causal queries—estimands that reference outcomes from different, mutually exclusive interventional worlds. Typical members are the probability of necessary and sufficient causation (PN, PS), treatment benefit and harm rates, the effect of persuasion, the distribution of the individual treatment effect (ITE), principal causal effects, principal fairness, and natural direct and indirect effects.

The key identification difficulty is twofold: beyond treatment–outcome confounding, one must characterize the dependence between potential outcomes, and randomization resolves only the former. This yields a sharp practical consequence: randomization alone does not identify any cross-world estimand, regardless of sample size or design quality. The paper reviews five strategies for the binary-treatment case:

  • Independence between potential outcomes, possibly latent conditional independence at the cost of parametric restrictions.
  • Monotonicity (Y(1)≥Y(0)Y(1) \geq Y(0) almost surely), which reduces the joint distribution to three unknown parameters identified by three moment equations. The paper highlights an asymmetry in its identifying power that practitioners should note: monotonicity suffices for point identification with multi-level ordered treatments and binary outcomes, but generally fails to point-identify joint distributions when outcomes are ordinal or continuous, even with binary treatment.
  • Association parameters (Pearson correlation or odds ratio) under ignorability, which turn identification into sensitivity analysis over the association parameter; positive association is argued to be plausible in applications and to substantially tighten bounds.
  • Copula models for continuous outcomes, since a scalar association parameter alone is insufficient.
  • Data fusion, where nonparametric identification of the joint distribution is achieved by combining multiple experimental studies.

Where these conditions are judged too restrictive, the paper points to partial identification via sharp bounds as the fallback, and it is candid that the required assumptions "may be too restrictive in practice."

Individual-level counterfactual queries: the second sublayer

Learning individual-level counterfactual outcomes is not a well-posed statistical task under a stochastic view, since Y(a)Y(a) is then a random variable rather than an estimand. The paper therefore introduces Pearl's deterministic viewpoint as an explicit assumption: each individual's potential outcome is a fixed quantity. This is a substantive commitment that departs from the Dawid-style decision-theoretic position, and the paper acknowledges the tension rather than resolving it.

The paper then illustrates why second-layer surrogates are dangerous for individual decisions. A subpopulation of ten individuals with CATE=0.1\text{CATE} = 0.1 may consist of five individuals with ITE =1= 1 and five with ITE =−0.8= -0.8; treating all ten harms half of them. This motivates complementing CATE-based rules with harm rates or ITE prediction intervals.

Three estimation strategies are reviewed. Pearl's abduction–action–prediction procedure computes counterfactuals by inferring exogenous variables from observed evidence, intervening in the SCM, and predicting. The paper is explicit about its two prerequisites—an SCM describing the data-generating process, and identifiability of exogenous variables (e.g., invertibility of the structural equation)—and states plainly that these may limit applicability. Rank preservation with quantile regression identifies counterfactual outcomes under ignorability plus strict monotonicity of the outcome in an exogenous variable, later generalized to rank preservation; the paper notes that this coupling coincides with the optimal transport map for one-dimensional distributions. Conformal inference constructs prediction intervals for the ITE with coverage guarantees under only ignorability and overlap, reducing to conformal prediction under covariate shift; intervals can be narrowed by imposing a positive cross-world association.

A comparison table in the paper makes the trade-off explicit: rank preservation yields point identification but requires stronger assumptions and does not generalize beyond populations with observed outcomes, whereas conformal inference is assumption-lean and generalizable but produces often-wide intervals with limited information about the ITE. Neither dominates.

Settings with post-treatment variables

The framework extends naturally to a post-treatment variable SS. Principal causal effects, E[Y(1)−Y(0)∣S(1)=s1,S(0)=s0]\mathbb{E}[Y(1) - Y(0) \mid S(1) = s_1, S(0) = s_0], are third-layer because they involve the joint distribution of (S(0),S(1))(S(0), S(1)), with applications to noncompliance, truncation by death, and surrogate evaluation. Short-term and long-term treatment effects remain second-layer, but long-term outcomes introduce a distinct obstacle—missingness from drop-out and limited follow-up—on top of confounding. In fairness, counterfactual parity is second-layer while Kusner-style counterfactual fairness and principal fairness are third-layer; the classification thus clarifies that these metrics differ not merely in formulation but in the strength of assumptions needed to identify them. In mediation, the total effect and controlled direct effect are second-layer, while natural direct and indirect effects are third-layer because terms like Y(1,M(0))Y(1, M(0)) are cross-world; this recovers the familiar result that NDE and NIE require assumptions beyond sequential ignorability.

Limitations and open questions

The paper is a position paper, and several boundaries are acknowledged. The classification is illustrative rather than exhaustive, and the treatment of the second layer is deliberately brief. The deterministic-viewpoint assumption underlying the individual-level sublayer is contested in the broader literature, and the paper does not adjudicate between stochastic and deterministic conceptions of counterfactuals. For cross-world estimands, all reviewed identification strategies impose conditions the paper itself concedes may be unrealistic, leaving partial identification as the only generally defensible route—an open question is which sensitivity parameterizations yield the tightest practically credible bounds. For individual-level inference, the trade-off between point identification under rank preservation and wide conformal intervals remains unresolved, and the paper does not offer criteria for choosing between them in a given application.

Conclusion

The paper supplies a formal potential outcomes criterion for assigning causal estimands to layers of Pearl's hierarchy, organized around marginal distributions, joint distributions and cross-world quantities, and individual-level counterfactuals. Its main contributions are the classification itself, the mapping of each class to its identification obstacles and available strategies, and the demonstration—via mediation, principal stratification, long-term effects, and fairness metrics—that the criterion cleanly separates second-layer from third-layer quantities that are easily conflated in practice. The framework's practical value lies in diagnosing whether a proposed estimand matches the scientific question and whether the accompanying assumptions are sufficient, insufficient, or unnecessarily strong.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 12 likes about this paper.