Papers
Topics
Authors
Recent
Search
2000 character limit reached

CTFIDU+ Algorithm for Counterfactual Identification

Updated 5 July 2026
  • The paper introduces ctfIDu⁺, an algorithm that extends counterfactual identification to handle arbitrary realizable Layer-3 distributions, and proves its soundness and completeness.
  • The method employs a two-layer approach with an inner routine (identify⁺) and an outer routine that factorizes counterfactual queries into ctf-factors over c-components in recursive SCMs.
  • Integrating counterfactual randomization (ctf-rand) allows the algorithm to tighten partial identification bounds and improve decision-making in non-parametric causal inference.

The CTFIDU+ algorithm—written in the source paper as ctfIDu+ctfIDu^+—is an identification procedure for counterfactual queries from an arbitrary collection of physically realizable input distributions, including observational, interventional, and certain counterfactual distributions obtainable via counterfactual randomization. It is introduced in “Causal Identification from Counterfactual Data: Completeness and Bounding Results” (Raghavan et al., 26 Feb 2026) to address a setting not handled by earlier identification algorithms: the input data themselves may belong to Layer 3 of the Pearl Causal Hierarchy, rather than only to observational or interventional layers. The algorithm is stated for unnested counterfactual queries over recursive structural causal models (SCMs), and the paper proves that it is sound and complete for this task (Raghavan et al., 26 Feb 2026).

1. Position within counterfactual identification

Counterfactual identification asks whether a target query P(Y=y)P(\mathbf{Y}_\star=\mathbf{y}) is uniquely computable from a causal graph G\mathcal{G} and a set of available distributions across all SCMs compatible with G\mathcal{G}. In the standard presentation of the Pearl Causal Hierarchy, observational distributions occupy L1\mathcal{L}_1, interventional distributions occupy L2\mathcal{L}_2, and counterfactual distributions occupy L3\mathcal{L}_3, including cross-world expressions such as

P(Yx=y,Zx=z,X=x).P\big(Y_x = y, Z_{x'} = z', X=x''\big).

Earlier completeness results were restricted to settings in which the input data lie in L1\mathcal{L}_1 or L2\mathcal{L}_2. IDC* is described as complete for P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})0 identification assuming full P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})1 data, while ctfID is described as complete for P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})2 identification assuming an arbitrary subset of P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})3 distributions. Other algorithms such as ID, IDC, and psIDC likewise assume that the inputs are observational or interventional rather than raw counterfactual distributions (Raghavan et al., 26 Feb 2026).

The motivation for P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})4 comes from counterfactual realizability, introduced by Raghavan and Bareinboim (2025). In that framework, some Layer-3 distributions can be sampled directly through experimental procedures. The key new experimental primitive is counterfactual randomization, denoted P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})5, which sets the value of P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})6 as perceived by a chosen child set P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})7 to a randomized value, while not overriding the unit’s natural value of P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})8 and not affecting the remaining children. This enlarges the set of admissible data-generating regimes beyond observation and standard randomized intervention, and yields a realizable sublayer P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})9 (Raghavan et al., 26 Feb 2026).

The resulting identification problem is therefore different from classical counterfactual identification: the input is no longer limited to observational and interventional distributions, but may be an arbitrary mix of realizable Layer-3 distributions. G\mathcal{G}0 is designed precisely for that generalized setting.

2. Formal setting and representation of queries

The algorithm is developed for recursive, acyclic SCMs G\mathcal{G}1, where G\mathcal{G}2 are observed variables, G\mathcal{G}3 are exogenous variables, and G\mathcal{G}4 are structural equations. Each model induces a semi-Markovian causal diagram with directed edges for observed parent relations and bidirected edges representing unmeasured confounding (Raghavan et al., 26 Feb 2026).

The target objects are general counterfactual events of the form G\mathcal{G}5, where G\mathcal{G}6 is a set of potential responses, possibly under different regimes, for example

G\mathcal{G}7

The corresponding Layer-3 distribution is G\mathcal{G}8. When all subscripts are identical, the expression collapses to a Layer-2 interventional distribution; when the subscripts are empty, it is observational (Raghavan et al., 26 Feb 2026).

The paper states G\mathcal{G}9 for unnested counterfactuals. Nested expressions are allowed in principle, but are to be handled by preprocessing, particularly the ancestral set transformation (AST). The assumptions listed for the algorithm are: an acyclic graph, discrete variables with finite domain, strict positivity of all distributions, and target queries that are unnested after preprocessing (Raghavan et al., 26 Feb 2026).

A central formal device is the ctf-factor, a Layer-3 generalization of Tian and Pearl’s c-factor. If

G\mathcal{G}0

then G\mathcal{G}1 denotes a counterfactual factor associated with a cluster of counterfactual variables sharing exogenous structure. The paper’s key structural claim is that any post-AST counterfactual distribution can be factorized into a product of ctf-factors over c-components, analogously to the factorization of interventional distributions. This factorization is the core reduction used by G\mathcal{G}2 (Raghavan et al., 26 Feb 2026).

The available input data are indexed by a set of action specifications

G\mathcal{G}3

where G\mathcal{G}4 represents observation, G\mathcal{G}5 represents standard intervention, and more general G\mathcal{G}6 may include G\mathcal{G}7. For each regime, a routine denoted G\mathcal{G}8 constructs the realizable counterfactual joint distribution observed under that regime (Raghavan et al., 26 Feb 2026).

3. Internal structure of G\mathcal{G}9

The algorithm has two layers. The inner routine, L1\mathcal{L}_10, attempts to identify one ctf-factor L1\mathcal{L}_11 from another ctf-factor L1\mathcal{L}_12. The outer routine, L1\mathcal{L}_13, reduces the full query to a collection of required ctf-factors and then searches the available regimes for input factors from which each required factor can be identified (Raghavan et al., 26 Feb 2026).

The L1\mathcal{L}_14 subroutine takes as input the graph L1\mathcal{L}_15, a target ctf-factor L1\mathcal{L}_16, and an available ctf-factor L1\mathcal{L}_17, subject to three conditions: L1\mathcal{L}_18, each observable appears at most once in L1\mathcal{L}_19, and L2\mathcal{L}_20 is a single c-component. It constructs a minimal closure L2\mathcal{L}_21 of L2\mathcal{L}_22 inside L2\mathcal{L}_23 such that no outside term appears in the subscripts of a term inside L2\mathcal{L}_24. If L2\mathcal{L}_25, identification is by marginalization: L2\mathcal{L}_26 If L2\mathcal{L}_27, the routine returns FAIL. Otherwise it marginalizes to L2\mathcal{L}_28, factorizes by c-components, selects the factor containing L2\mathcal{L}_29, and recurses (Raghavan et al., 26 Feb 2026).

The outer algorithm proceeds by normalizing the target query, factorizing it, matching the resulting factors against available experimental regimes, and assembling the final formula. Its high-level workflow is as follows.

Step Operation Result
1 Simplify redundant subscripts Remove redundant or inconsistent assignments
2 Apply AST to ancestors Rewrite target as a marginal of L3\mathcal{L}_30
3 Factorize into ctf-factors Obtain necessary and sufficient factors L3\mathcal{L}_31
4 Process each input regime Construct and factorize realizable L3\mathcal{L}_32
5 Run L3\mathcal{L}_33 factorwise Express each target factor from some input factor
6 Assemble or fail Return final formula or FAIL

The algorithm first simplifies the target using an exclusion lemma. If conflicting assignments occur, such as two terms L3\mathcal{L}_34 and L3\mathcal{L}_35 with L3\mathcal{L}_36, it returns L3\mathcal{L}_37, corresponding to an impossible event. It then computes the ancestor expansion L3\mathcal{L}_38 and applies AST to rewrite

L3\mathcal{L}_39

Next, it partitions P(Yx=y,Zx=z,X=x).P\big(Y_x = y, Z_{x'} = z', X=x''\big).0 into P(Yx=y,Zx=z,X=x).P\big(Y_x = y, Z_{x'} = z', X=x''\big).1 such that each P(Yx=y,Zx=z,X=x).P\big(Y_x = y, Z_{x'} = z', X=x''\big).2 is a c-component in the induced graph, yielding the factorization

P(Yx=y,Zx=z,X=x).P\big(Y_x = y, Z_{x'} = z', X=x''\big).3

These are described as the necessary and sufficient ctf-factors for identifying the query (Raghavan et al., 26 Feb 2026).

For each input regime P(Yx=y,Zx=z,X=x).P\big(Y_x = y, Z_{x'} = z', X=x''\big).4, the algorithm constructs the corresponding realizable distribution P(Yx=y,Zx=z,X=x).P\big(Y_x = y, Z_{x'} = z', X=x''\big).5, applies AST, partitions into c-components, and factorizes into ctf-factors P(Yx=y,Zx=z,X=x).P\big(Y_x = y, Z_{x'} = z', X=x''\big).6. It then searches across all regimes and all such factors to find, for each target factor P(Yx=y,Zx=z,X=x).P\big(Y_x = y, Z_{x'} = z', X=x''\big).7, an input factor P(Yx=y,Zx=z,X=x).P\big(Y_x = y, Z_{x'} = z', X=x''\big).8 with P(Yx=y,Zx=z,X=x).P\big(Y_x = y, Z_{x'} = z', X=x''\big).9 for which L1\mathcal{L}_10 succeeds. If every target factor is identified, the returned formula is

L1\mathcal{L}_11

If any required factor cannot be recovered from any regime, the output is FAIL (Raghavan et al., 26 Feb 2026).

4. Obstructions, completeness, and relation to earlier algorithms

The negative structure underlying failure is the ctf-hedge, defined through a more primitive object called a ctf-forest. A ctf-forest is a collection L1\mathcal{L}_12 satisfying four properties: each observed variable appears at most once; the induced subgraph on L1\mathcal{L}_13 is a c-component whose bidirected edges form a minimum spanning tree; L1\mathcal{L}_14 for some L1\mathcal{L}_15; and each vertex has at most one child. A ctf-hedge is a ctf-forest rooted at L1\mathcal{L}_16 that strictly contains the root set and satisfies an additional “value chain” condition tying parent values to child subscripts (Raghavan et al., 26 Feb 2026).

The paper states a non-identifiability lemma: if L1\mathcal{L}_17 is a ctf-hedge rooted at L1\mathcal{L}_18, then L1\mathcal{L}_19 is not identifiable from L2\mathcal{L}_20 given L2\mathcal{L}_21. The proof idea uses two SCMs with the same minimum spanning tree but different value assignments encoded in a bit-representation of edges; the models agree on L2\mathcal{L}_22 but disagree on L2\mathcal{L}_23 (Raghavan et al., 26 Feb 2026).

This obstruction yields the characterization of the inner routine: for suitable input factors,

L2\mathcal{L}_24

The same logic is lifted to the outer routine in Theorem 4.3, which states that for an unnested counterfactual expression L2\mathcal{L}_25, the query L2\mathcal{L}_26 is identifiable from L2\mathcal{L}_27 and regime set L2\mathcal{L}_28 if and only if L2\mathcal{L}_29 returns an expression (Raghavan et al., 26 Feb 2026).

The soundness argument relies on the correctness of AST, ctf-factorization, marginalization, and the soundness of P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})00. The completeness argument relies on the fact that AST and factorization isolate a minimal collection of necessary ctf-factors; if one of those factors cannot be identified from any regime, then the corresponding failure of P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})01 implies a ctf-hedge obstruction, and therefore there exist two SCMs agreeing on all inputs in P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})02 but disagreeing on the target (Raghavan et al., 26 Feb 2026).

The paper positions P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})03 as a strict generalization of prior identification algorithms. When P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})04 contains only observational and interventional regimes, it reduces to ctfID, and in the special case of full P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})05 availability it reduces to IDC*. When P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})06 includes realizable Layer-3 distributions produced via ctf-rand, it extends beyond the scope of ID, IDC, IDC*, and ctfID by directly using counterfactual data as input (Raghavan et al., 26 Feb 2026).

A common misconception addressed by this development is that Layer-3 distributions are necessarily inaccessible except through formal identification. The paper’s framework rejects that blanket assumption: some Layer-3 distributions are realizable, but not all of them. That distinction is essential to both the power and the limits of P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})07.

5. Realizability, hierarchy refinements, and the limit of exact identification

The paper refines the causal hierarchy beyond P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})08 by introducing intermediate realizable sublayers. P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})09 is a subclass of realizable counterfactual distributions in which ctf-rand cannot be path-specific. P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})10 is the set of all distributions realizable in principle by allowing path-specific ctf-rand on every edge in the graph. The complement P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})11 consists of purely counterfactual distributions that remain unrealizable even under maximal ctf-rand capability (Raghavan et al., 26 Feb 2026).

Using P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})12, the paper proves a limit of identification theorem. If a query P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})13 belongs to layer P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})14 but not to any lower layer, then for any P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})15, there exists a graph in which P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})16 is identifiable from data in layer P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})17, except when P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})18. The critical consequence is that there are no purely Layer-3 queries in P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})19 that are identifiable from P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})20 data (Raghavan et al., 26 Feb 2026).

This establishes P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})21 as the paper’s theoretical limit of exact causal identification in the non-parametric setting. The source further presents an informal corollary described as an identifiability–realizability duality: a query is identifiable from observational plus experimental data and P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})22 if and only if it is realizable in principle via ctf-rand actions (Raghavan et al., 26 Feb 2026).

The implication is not that every counterfactual quantity becomes directly measurable, but rather that the boundary of exact identification coincides with the boundary of physical realizability under the allowed experimental primitives. A plausible implication is that, within this framework, advances in counterfactual experimentation enlarge exact identifiability only insofar as they enlarge the realizable sublayer itself.

The paper gives a concrete class of exceptions: queries such as P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})23 may lie in P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})24, and are therefore non-identifiable from any realizable counterfactual data. It also states that the natural total effect (NTE) of Leek et al. (2025), used in XAI, depends on probabilities of causation of the form

P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})25

which likewise lie in P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})26, and are therefore not exactly identifiable even under maximal ctf-randomization (Raghavan et al., 26 Feb 2026).

6. Partial identification, analytic bounds, and simulation findings

Because some counterfactual queries are irreducibly non-identifiable, the paper turns to partial identification. Given a graph P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})27, a target non-identifiable query P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})28, and data regimes P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})29, the task is to characterize the tight range P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})30 over all SCMs consistent with the graph and the observed data constraints. The paper states a monotonicity proposition: if P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})31, then the tight bounds under P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})32 are contained in those under P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})33 (Raghavan et al., 26 Feb 2026).

The analytic illustration is the bow graph, with P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})34 and an unobserved confounder between P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})35 and P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})36. In that model, the NTE reduces to a function involving the probability of causation P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})37. Three data scenarios are compared (Raghavan et al., 26 Feb 2026).

With only observational data P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})38, the paper states that the tight bounds are completely uninformative: P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})39

With observational plus interventional data, specifically P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})40 and P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})41 for each P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})42, the paper gives Balke–Pearl-style linear-programming bounds: P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})43 where

P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})44

When realizable counterfactual data are added, specifically P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})45 for all P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})46, the paper gives sharper bounds: P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})47 together with the guarantee

P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})48

Thus the added Layer-P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})49 data systematically tighten the bounds for this non-identifiable quantity (Raghavan et al., 26 Feb 2026).

The simulation section uses a Bayesian sampling scheme, pID after Zhang et al. (2022), to estimate credible intervals from finite synthetic samples. In the “Traffic Camera – version 2” example, the graph includes confounding between both P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})50 and P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})51, with P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})52 and P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})53. The queries are the natural direct effect P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})54 and the NTE component P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})55. The reported findings are that, for NTE, the credible interval under P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})56 is significantly narrower than under P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})57, and that for NDE the interval collapses to the true value once counterfactual data are included, consistent with exact identifiability under ctf-rand (Raghavan et al., 26 Feb 2026).

In the “Unit Selection for treatment assignment” example, the same bow graph is used for de-addiction treatment. The paper distinguishes four unit types—Always-0, Helped, Hurt, and Always-1—with corresponding potential outcomes and benefits. It compares a standard interventional strategy, based on P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})58 bounds on P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})59, with a counterfactual strategy using ctf-rand to estimate

P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})60

and thereby conditional benefits P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})61. The reported finding is that the counterfactual strategy yields positive bounds for the P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})62 subpopulation and negative bounds for the P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})63 subpopulation, implying an optimal policy that treats only units with natural P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})64, a policy stated to be unattainable using P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})65 alone (Raghavan et al., 26 Feb 2026).

These results support two distinct conclusions. First, realizable counterfactual data can convert some previously non-identifiable quantities into exactly identifiable ones. Second, when exact identification remains impossible, the same data can materially sharpen partial identification and enable more refined decision rules. This suggests that the main significance of P(Y=y)P(\mathbf{Y}_\star=\mathbf{y})66 is not merely algorithmic unification, but the precise delineation of what counterfactual experimentation can and cannot buy in non-parametric causal inference (Raghavan et al., 26 Feb 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to CTFIDU+ Algorithm.