---
title: CTFIDU+ Algorithm for Counterfactual Identification
url: https://www.emergentmind.com/topics/ctfidu-algorithm
type: topic
---

# CTFIDU+ Algorithm for Counterfactual Identification

The **CTFIDU+ algorithm**—written in the source paper as \(ctfIDu^+\)—is an identification procedure for counterfactual queries from an arbitrary collection of **physically realizable** input distributions, including observational, interventional, and certain counterfactual distributions obtainable via **counterfactual randomization**. It is introduced in “Causal Identification from Counterfactual Data: Completeness and Bounding Results” [2602.23541] to address a setting not handled by earlier identification algorithms: the input data themselves may belong to Layer 3 of the Pearl Causal Hierarchy, rather than only to observational or interventional layers. The algorithm is stated for **unnested** counterfactual queries over recursive structural causal models (SCMs), and the paper proves that it is **sound and complete** for this task [2602.23541].

## 1. Position within counterfactual identification

Counterfactual identification asks whether a target query \(P(\mathbf{Y}_\star=\mathbf{y})\) is uniquely computable from a causal graph \(\mathcal{G}\) and a set of available distributions across all SCMs compatible with \(\mathcal{G}\). In the standard presentation of the Pearl Causal Hierarchy, observational distributions occupy \(\mathcal{L}_1\), interventional distributions occupy \(\mathcal{L}_2\), and counterfactual distributions occupy \(\mathcal{L}_3\), including cross-world expressions such as
\[
P\big(Y_x = y, Z_{x'} = z', X=x''\big).
\]

Earlier completeness results were restricted to settings in which the input data lie in \(\mathcal{L}_1\) or \(\mathcal{L}_2\). IDC* is described as complete for \(\mathcal{L}_3\) identification assuming full \(\mathcal{L}_2\) data, while ctfID is described as complete for \(\mathcal{L}_3\) identification assuming an arbitrary subset of \(\mathcal{L}_2\) distributions. Other algorithms such as ID, IDC, and psIDC likewise assume that the inputs are observational or interventional rather than raw counterfactual distributions [2602.23541].

The motivation for \(ctfIDu^+\) comes from **counterfactual realizability**, introduced by Raghavan and Bareinboim (2025). In that framework, some Layer-3 distributions can be sampled directly through experimental procedures. The key new experimental primitive is **counterfactual randomization**, denoted \(\text{ctf-rand}(X\to\mathbf{C})\), which sets the value of \(X\) as perceived by a chosen child set \(\mathbf{C}\subseteq \mathrm{Ch}(X)\) to a randomized value, while not overriding the unit’s natural value of \(X\) and not affecting the remaining children. This enlarges the set of admissible data-generating regimes beyond observation and standard randomized intervention, and yields a realizable sublayer \(\mathcal{L}_{2.5}\subseteq \mathcal{L}_3\) [2602.23541].

The resulting identification problem is therefore different from classical counterfactual identification: the input is no longer limited to observational and interventional distributions, but may be an **arbitrary mix of realizable Layer-3 distributions**. \(ctfIDu^+\) is designed precisely for that generalized setting.

## 2. Formal setting and representation of queries

The algorithm is developed for recursive, acyclic SCMs \(\mathcal{M}=\langle \mathbf{V},\mathbf{U},\mathcal{F},P(\mathbf{U})\rangle\), where \(\mathbf{V}\) are observed variables, \(\mathbf{U}\) are exogenous variables, and \(\mathcal{F}=\{f_i\}\) are structural equations. Each model induces a semi-Markovian causal diagram with directed edges for observed parent relations and bidirected edges representing unmeasured confounding [2602.23541].

The target objects are general counterfactual events of the form \((\mathbf{W}_\star=\mathbf{w})\), where \(\mathbf{W}_\star\) is a set of potential responses, possibly under different regimes, for example
\[
(Y_x=y,\;Y_{x'}=y',\;X=x'').
\]
The corresponding Layer-3 distribution is \(P(\mathbf{W}_\star=\mathbf{w})\). When all subscripts are identical, the expression collapses to a Layer-2 interventional distribution; when the subscripts are empty, it is observational [2602.23541].

The paper states \(ctfIDu^+\) for **unnested** counterfactuals. Nested expressions are allowed in principle, but are to be handled by preprocessing, particularly the **ancestral set transformation** (AST). The assumptions listed for the algorithm are: an acyclic graph, discrete variables with finite domain, strict positivity of all distributions, and target queries that are unnested after preprocessing [2602.23541].

A central formal device is the **ctf-factor**, a Layer-3 generalization of Tian and Pearl’s c-factor. If
\[
\mathbf{C}_\star=\{V_{1[\mathbf{pa}_1]},\dots,V_{k[\mathbf{pa}_k]}\},
\]
then \(Q[\mathbf{C}_\star](\mathbf{c})\) denotes a counterfactual factor associated with a cluster of counterfactual variables sharing exogenous structure. The paper’s key structural claim is that any post-AST counterfactual distribution can be factorized into a product of ctf-factors over c-components, analogously to the factorization of interventional distributions. This factorization is the core reduction used by \(ctfIDu^+\) [2602.23541].

The available input data are indexed by a set of action specifications
\[
\mathbb{A}=\{\mathcal{A}_1,\dots,\mathcal{A}_k\},
\]
where \(\mathcal{A}=\emptyset\) represents observation, \(\mathcal{A}=\{\text{rand}(\mathbf{X})\}\) represents standard intervention, and more general \(\mathcal{A}\) may include \(\text{ctf-rand}(X\to\mathbf{C})\). For each regime, a routine denoted \(regime\text{-}regex(\mathcal{G},\mathcal{A})\) constructs the realizable counterfactual joint distribution observed under that regime [2602.23541].

## 3. Internal structure of \(ctfIDu^+\)

The algorithm has two layers. The inner routine, \(identify^+\), attempts to identify one ctf-factor \(Q[\mathbf{C}_\star](\mathbf{c})\) from another ctf-factor \(Q[\mathbf{T}_\star](\mathbf{t})\). The outer routine, \(ctfIDu^+\), reduces the full query to a collection of required ctf-factors and then searches the available regimes for input factors from which each required factor can be identified [2602.23541].

The \(identify^+\) subroutine takes as input the graph \(\mathcal{G}\), a target ctf-factor \(Q[\mathbf{C}_\star](\mathbf{c})\), and an available ctf-factor \(Q[\mathbf{T}_\star](\mathbf{t})\), subject to three conditions: \(\mathbf{C}_\star\subseteq \mathbf{T}_\star\), each observable appears at most once in \(\mathbf{T}_\star\), and \(\mathcal{G}[\mathbf{V}(\mathbf{T}_\star)]\) is a single c-component. It constructs a minimal closure \(\mathbf{H}_\star\) of \(\mathbf{C}_\star\) inside \(\mathbf{T}_\star\) such that no outside term appears in the subscripts of a term inside \(\mathbf{H}_\star\). If \(\mathbf{H}_\star=\mathbf{C}_\star\), identification is by marginalization:
\[
Q[\mathbf{C}_\star](\mathbf{c})=\sum_{\mathbf{t}\setminus \mathbf{c}}Q[\mathbf{T}_\star](\mathbf{t}).
\]
If \(\mathbf{H}_\star=\mathbf{T}_\star\), the routine returns FAIL. Otherwise it marginalizes to \(Q[\mathbf{H}_\star]\), factorizes by c-components, selects the factor containing \(\mathbf{C}_\star\), and recurses [2602.23541].

The outer algorithm proceeds by normalizing the target query, factorizing it, matching the resulting factors against available experimental regimes, and assembling the final formula. Its high-level workflow is as follows.

| Step | Operation | Result |
|---|---|---|
| 1 | Simplify redundant subscripts | Remove redundant or inconsistent assignments |
| 2 | Apply AST to ancestors | Rewrite target as a marginal of \(P(\mathbf{W}_\star=\mathbf{w})\) |
| 3 | Factorize into ctf-factors | Obtain necessary and sufficient factors \(Q[\mathbf{C}^j_\star]\) |
| 4 | Process each input regime | Construct and factorize realizable \(P(\mathbf{T}_\star=\mathbf{t})\) |
| 5 | Run \(identify^+\) factorwise | Express each target factor from some input factor |
| 6 | Assemble or fail | Return final formula or FAIL |

The algorithm first simplifies the target using an exclusion lemma. If conflicting assignments occur, such as two terms \(y_x\) and \(y'_x\) with \(y\neq y'\), it returns \(0\), corresponding to an impossible event. It then computes the ancestor expansion \(\mathbf{W}_\star=An(\mathbf{Y}_\star)\) and applies AST to rewrite
\[
P(\mathbf{Y}_\star=\mathbf{y})
=
\sum_{\mathbf{w}\setminus \mathbf{y}}P(\mathbf{W}_\star=\mathbf{w}).
\]
Next, it partitions \(\mathbf{W}_\star\) into \(\mathbf{C}^1_\star,\dots,\mathbf{C}^k_\star\) such that each \(\mathbf{V}(\mathbf{C}^j_\star)\) is a c-component in the induced graph, yielding the factorization
\[
P(\mathbf{W}_\star=\mathbf{w})=\prod_j Q[\mathbf{C}^j_\star](\mathbf{c}^j).
\]
These are described as the **necessary and sufficient ctf-factors** for identifying the query [2602.23541].

For each input regime \(\mathcal{A}\in\mathbb{A}\), the algorithm constructs the corresponding realizable distribution \(P(\mathbf{T}_\star=\mathbf{t})\), applies AST, partitions into c-components, and factorizes into ctf-factors \(Q[\mathbf{T}^i_\star](\mathbf{t}^i)\). It then searches across all regimes and all such factors to find, for each target factor \(Q[\mathbf{C}^j_\star]\), an input factor \(Q[\mathbf{T}^i_\star]\) with \(\mathbf{C}^j_\star\subseteq \mathbf{T}^i_\star\) for which \(identify^+\) succeeds. If every target factor is identified, the returned formula is
\[
P(\mathbf{Y}_\star=\mathbf{y})
=
\sum_{\mathbf{w}\setminus \mathbf{y}} \prod_j Q[\mathbf{C}^j_\star](\mathbf{c}^j).
\]
If any required factor cannot be recovered from any regime, the output is FAIL [2602.23541].

## 4. Obstructions, completeness, and relation to earlier algorithms

The negative structure underlying failure is the **ctf-hedge**, defined through a more primitive object called a **ctf-forest**. A ctf-forest is a collection \(\{\mathbf{T}_\star=\mathbf{t}\}\) satisfying four properties: each observed variable appears at most once; the induced subgraph on \(\mathbf{T}=\mathbf{V}(\mathbf{T}_\star)\) is a c-component whose bidirected edges form a minimum spanning tree; \(\mathbf{T}=An(\mathbf{C})_{\mathcal{G}[\mathbf{T}]}\) for some \(\mathbf{C}_\star\subseteq \mathbf{T}_\star\); and each vertex has at most one child. A ctf-hedge is a ctf-forest rooted at \(\mathbf{C}_\star\) that strictly contains the root set and satisfies an additional “value chain” condition tying parent values to child subscripts [2602.23541].

The paper states a non-identifiability lemma: if \(\{\mathbf{T}_\star=\mathbf{t}\}\) is a ctf-hedge rooted at \(\{\mathbf{C}_\star=\mathbf{c}\}\), then \(Q[\mathbf{C}_\star](\mathbf{c})\) is not identifiable from \(Q[\mathbf{T}_\star](\mathbf{t})\) given \(\mathcal{G}\). The proof idea uses two SCMs with the same minimum spanning tree but different value assignments encoded in a bit-representation of edges; the models agree on \(P(\mathbf{T}_\star)\) but disagree on \(P(\mathbf{C}_\star)\) [2602.23541].

This obstruction yields the characterization of the inner routine: for suitable input factors,
\[
Q[\mathbf{C}_\star](\mathbf{c}) \text{ is identifiable from } Q[\mathbf{T}_\star](\mathbf{t})
\iff identify^+ \text{ returns an expression.}
\]
The same logic is lifted to the outer routine in Theorem 4.3, which states that for an unnested counterfactual expression \(\mathbf{Y}_\star\), the query \(P(\mathbf{Y}_\star=\mathbf{y})\) is identifiable from \(\mathcal{G}\) and regime set \(\mathbb{A}\) **if and only if** \(ctfIDu^+\) returns an expression [2602.23541].

The soundness argument relies on the correctness of AST, ctf-factorization, marginalization, and the soundness of \(identify^+\). The completeness argument relies on the fact that AST and factorization isolate a minimal collection of necessary ctf-factors; if one of those factors cannot be identified from any regime, then the corresponding failure of \(identify^+\) implies a ctf-hedge obstruction, and therefore there exist two SCMs agreeing on all inputs in \(\mathbb{A}\) but disagreeing on the target [2602.23541].

The paper positions \(ctfIDu^+\) as a strict generalization of prior identification algorithms. When \(\mathbb{A}\) contains only observational and interventional regimes, it reduces to ctfID, and in the special case of full \(\mathcal{L}_2\) availability it reduces to IDC*. When \(\mathbb{A}\) includes realizable Layer-3 distributions produced via ctf-rand, it extends beyond the scope of ID, IDC, IDC*, and ctfID by directly using counterfactual data as input [2602.23541].

A common misconception addressed by this development is that Layer-3 distributions are necessarily inaccessible except through formal identification. The paper’s framework rejects that blanket assumption: some Layer-3 distributions are realizable, but not all of them. That distinction is essential to both the power and the limits of \(ctfIDu^+\).

## 5. Realizability, hierarchy refinements, and the limit of exact identification

The paper refines the causal hierarchy beyond \(\mathcal{L}_1,\mathcal{L}_2,\mathcal{L}_3\) by introducing intermediate realizable sublayers. \(\mathcal{L}_{2.25}\) is a subclass of realizable counterfactual distributions in which ctf-rand cannot be path-specific. \(\mathcal{L}_{2.5}\) is the set of all distributions realizable in principle by allowing path-specific ctf-rand on every edge in the graph. The complement \(\mathcal{L}_3\setminus\mathcal{L}_{2.5}\) consists of purely counterfactual distributions that remain unrealizable even under maximal ctf-rand capability [2602.23541].

Using \(ctfIDu^+\), the paper proves a **limit of identification** theorem. If a query \(Q\) belongs to layer \(\mathcal{L}_i\) but not to any lower layer, then for any \(j<i\), there exists a graph in which \(Q\) is identifiable from data in layer \(\mathcal{L}_j\), **except when \(i=3\)**. The critical consequence is that there are **no** purely Layer-3 queries in \(\mathcal{L}_3\setminus\mathcal{L}_{2.5}\) that are identifiable from \(\mathcal{L}_{2.5}\) data [2602.23541].

This establishes \(\mathcal{L}_{2.5}\) as the paper’s **theoretical limit of exact causal identification in the non-parametric setting**. The source further presents an informal corollary described as an **identifiability–realizability duality**: a query is identifiable from observational plus experimental data and \(\mathcal{G}\) if and only if it is realizable in principle via ctf-rand actions [2602.23541].

The implication is not that every counterfactual quantity becomes directly measurable, but rather that the boundary of exact identification coincides with the boundary of physical realizability under the allowed experimental primitives. A plausible implication is that, within this framework, advances in counterfactual experimentation enlarge exact identifiability only insofar as they enlarge the realizable sublayer itself.

The paper gives a concrete class of exceptions: queries such as \(P(y_x\mid z_{x'},x'')\) may lie in \(\mathcal{L}_3\setminus\mathcal{L}_{2.5}\), and are therefore non-identifiable from any realizable counterfactual data. It also states that the **natural total effect (NTE)** of Leek et al. (2025), used in XAI, depends on probabilities of causation of the form
\[
P(y_x\mid x',y'),
\]
which likewise lie in \(\mathcal{L}_3\setminus\mathcal{L}_{2.5}\), and are therefore not exactly identifiable even under maximal ctf-randomization [2602.23541].

## 6. Partial identification, analytic bounds, and simulation findings

Because some counterfactual queries are irreducibly non-identifiable, the paper turns to **partial identification**. Given a graph \(\mathcal{G}\), a target non-identifiable query \(Q=P(\mathbf{y}_\star)\), and data regimes \(\mathbb{A}\), the task is to characterize the tight range \([l,r]\subseteq[0,1]\) over all SCMs consistent with the graph and the observed data constraints. The paper states a monotonicity proposition: if \(\mathbb{A}'\supset \mathbb{A}\), then the tight bounds under \(\mathbb{A}'\) are contained in those under \(\mathbb{A}\) [2602.23541].

The analytic illustration is the **bow graph**, with \(X\to Y\) and an unobserved confounder between \(X\) and \(Y\). In that model, the NTE reduces to a function involving the probability of causation \(P(y_x\mid x',y')\). Three data scenarios are compared [2602.23541].

With only observational data \(P(X,Y)\), the paper states that the tight bounds are completely uninformative:
\[
P(y_x\mid x',y'),\quad x\neq x' \in [0,1].
\]

With observational plus interventional data, specifically \(P(X,Y)\) and \(P(Y_x)\) for each \(x\), the paper gives Balke–Pearl-style linear-programming bounds:
\[
\begin{aligned}
l &= \max\left\{0,\frac{\alpha_{\min}-(1-P(y' \mid x'))}{P(y' \mid x')}\right\}, \\
r &= \min\left\{1,\frac{\alpha_{\max}}{P(y' \mid x')}\right\},
\end{aligned}
\]
where
\[
\begin{aligned}
\alpha_{\min} &= \max\left\{0,\frac{P(y_x)-(1-P(x'))}{P(x')}\right\},\\
\alpha_{\max} &= \min\left\{1,\frac{P(y_x)}{P(x')}\right\}.
\end{aligned}
\]

When realizable counterfactual data are added, specifically \(P(Y_x\mid X)\) for all \(x\), the paper gives sharper bounds:
\[
\begin{aligned}
l' &= \max\left\{0,\frac{P(y_x\mid x')-(1-P(y' \mid x'))}{P(y' \mid x')}\right\},\\
r' &= \min\left\{1,\frac{P(y_x\mid x')}{P(y' \mid x')}\right\},
\end{aligned}
\]
together with the guarantee
\[
[l',r']\subseteq [l,r].
\]
Thus the added Layer-\(2.5\) data systematically tighten the bounds for this non-identifiable quantity [2602.23541].

The simulation section uses a Bayesian sampling scheme, pID after Zhang et al. (2022), to estimate credible intervals from finite synthetic samples. In the “Traffic Camera – version 2” example, the graph includes confounding between both \((X,Y)\) and \((Z,Y)\), with \(X\in\{0,1,2\}\) and \(Y,Z\in\{0,1\}\). The queries are the **natural direct effect** \(P(Y_{xZ_{x'}})\) and the NTE component \(P(y_x\mid x',y')\). The reported findings are that, for NTE, the credible interval under \(\mathcal{L}_{2.5}\) is significantly narrower than under \(\mathcal{L}_2\), and that for NDE the interval collapses to the true value once counterfactual data are included, consistent with exact identifiability under ctf-rand [2602.23541].

In the “Unit Selection for treatment assignment” example, the same bow graph is used for de-addiction treatment. The paper distinguishes four unit types—Always-0, Helped, Hurt, and Always-1—with corresponding potential outcomes and benefits. It compares a standard interventional strategy, based on \(\mathcal{L}_1+\mathcal{L}_2\) bounds on \(P(y'_0,y_1)\), with a counterfactual strategy using ctf-rand to estimate
\[
P(y'_0,y_1\mid X=x'),\quad x'\in\{0,1\},
\]
and thereby conditional benefits \(\Delta(1\mid X=x')\). The reported finding is that the counterfactual strategy yields positive bounds for the \(X=0\) subpopulation and negative bounds for the \(X=1\) subpopulation, implying an optimal policy that treats only units with natural \(X=0\), a policy stated to be unattainable using \(\mathcal{L}_2\) alone [2602.23541].

These results support two distinct conclusions. First, realizable counterfactual data can convert some previously non-identifiable quantities into exactly identifiable ones. Second, when exact identification remains impossible, the same data can materially sharpen partial identification and enable more refined decision rules. This suggests that the main significance of \(ctfIDu^+\) is not merely algorithmic unification, but the precise delineation of what counterfactual experimentation can and cannot buy in non-parametric causal inference [2602.23541].

Source: https://www.emergentmind.com/topics/ctfidu-algorithm