Papers
Topics
Authors
Recent
Search
2000 character limit reached

The Causal Description Gap: Information-Theoretic Separations Across Pearl's Hierarchy

Published 4 May 2026 in stat.ML, cs.AI, cs.IT, and cs.LG | (2605.02177v1)

Abstract: Pearl's causal hierarchy shows that observational, interventional, and counterfactual queries are qualitatively distinct. We ask a quantitative version of this question: how many additional bits are needed to specify higher-rung causal answers once lower-rung answers are known? We formalize this via query-class description length, the Kolmogorov complexity of the answer oracle induced by an SCM for a class of queries. Our main construction gives binary acyclic SCMs whose observational distribution has constant description length, while the single-variable interventional answer oracle has description length Θ(n<sup>2)Θ(n<sup>2). A degree-sensitive upper bound shows that finite-gate-schema SCMs of indegree dd have observational-interventional gap at most O(ndlog(en/d)+nlogn)O(nd \log(en/d) + n \log n), making the quadratic construction order-optimal in the dense regime and a rooted-tree construction order-optimal for bounded indegree. The quadratic separation persists under ε\varepsilon-accurate total-variation descriptions for every fixed $\varepsilon &lt; 1/4$. At the next rung, the full hard-do interventional oracle can still leave a Θ(n)Θ(n) counterfactual description gap. A general ambiguity-to-bits theorem and Shannon analogue show that these gaps equal the logarithm of residual higher-rung ambiguity up to lower-order terms.

Authors (1)

Summary

  • The paper formalizes causal non-identifiability as a description-length gap and shows that binary SCMs with constant-complexity observations can require Θ(n²) bits to specify their interventional oracles.
  • Degree-sensitive bounds show bounded-indegree SCMs have gaps of at most O(n log n), while dense mechanisms achieve the quadratic limit, with results remaining valid under fixed total-variation error.
  • The paper also demonstrates a Θ(n) counterfactual gap after complete hard interventions and proves observational data cannot identify hidden mechanisms when all mechanisms share the same distribution.

Overview

The paper under review, "The Causal Description Gap: Information-Theoretic Separations Across Pearl's Hierarchy" (2605.02177), converts a qualitative fact about causal inference into a quantitative one. The Causal Hierarchy Theorem establishes that observational (rung 1), interventional (rung 2), and counterfactual (rung 3) queries are generically non-identifiable from lower-rung information [bareinboim2022pearl]. The author asks: when identification fails, how many bits of residual information remain? The answer is formalized through query-class description length — the Kolmogorov complexity K(AnsQ(M)n)K(Ans_Q(M) \mid n) of the answer oracle induced by a structural causal model (SCM) for a query class QQ — and the causal description gap Δ21(M)=K(Int1(M)Obs(M),n)\Delta_{2|1}(M) = K(Int_1(M) \mid Obs(M), n), the additional bits needed to specify the interventional oracle given the observational one.

The headline result is a family of binary acyclic SCMs whose observational distribution has constant description length while the single-variable interventional oracle requires Θ(n2)\Theta(n^2) bits. This refines non-identifiability from a Boolean obstruction into a scaling information measure: the familiar two-variable ambiguity (XYX \to Y versus XYX \leftarrow Y via deterministic copy) is scaled to 2Θ(n2)2^{\Theta(n^2)} residual mechanisms.

Framework

For an SCM MM on binary variables, three canonical lengths are defined: DL1(M):=K(Obs(M)n)DL_1(M) := K(Obs(M) \mid n), DL2(M):=K(Int1(M)n)DL_2(M) := K(Int_1(M) \mid n) where QQ0 comprises the observational distribution plus all QQ1 single-variable interventional distributions, and QQ2 for single-node counterfactual oracles over parallel worlds sharing exogenous noise. Distributions are encoded as strings of rational probabilities; encoding choices affect results only by additive constants. A counting lemma supplies the workhorse lower bound: among QQ3 distinct strings, at least one has conditional complexity at least QQ4, and all but a QQ5 fraction exceed QQ6. Since every construction here has QQ7, a simple reduction shows QQ8, so gap and length coincide up to lower-order terms throughout.

The quadratic separation

The central construction hides an arbitrary bipartite graph QQ9 with Δ21(M)=K(Int1(M)Obs(M),n)\Delta_{2|1}(M) = K(Int_1(M) \mid Obs(M), n)0, Δ21(M)=K(Int1(M)Obs(M),n)\Delta_{2|1}(M) = K(Int_1(M) \mid Obs(M), n)1, inside an SCM. A root Δ21(M)=K(Int1(M)Obs(M),n)\Delta_{2|1}(M) = K(Int_1(M) \mid Obs(M), n)2 feeds layer Δ21(M)=K(Int1(M)Obs(M),n)\Delta_{2|1}(M) = K(Int_1(M) \mid Obs(M), n)3 by copy gates; each layer-Δ21(M)=K(Int1(M)Obs(M),n)\Delta_{2|1}(M) = K(Int_1(M) \mid Obs(M), n)4 node is an AND of Δ21(M)=K(Int1(M)Obs(M),n)\Delta_{2|1}(M) = K(Int_1(M) \mid Obs(M), n)5 and its Δ21(M)=K(Int1(M)Obs(M),n)\Delta_{2|1}(M) = K(Int_1(M) \mid Obs(M), n)6-neighbors in Δ21(M)=K(Int1(M)Obs(M),n)\Delta_{2|1}(M) = K(Int_1(M) \mid Obs(M), n)7. Observationally, every choice of Δ21(M)=K(Int1(M)Obs(M),n)\Delta_{2|1}(M) = K(Int_1(M) \mid Obs(M), n)8 yields the identical two-point distribution on Δ21(M)=K(Int1(M)Obs(M),n)\Delta_{2|1}(M) = K(Int_1(M) \mid Obs(M), n)9, so Θ(n2)\Theta(n^2)0. But Θ(n2)\Theta(n^2)1 forces neighbors of Θ(n2)\Theta(n^2)2 deterministically to 0 while non-neighbors remain uniform, so the Θ(n2)\Theta(n^2)3 row-interventions decode the full adjacency matrix. With Θ(n2)\Theta(n^2)4 distinct graphs mapping injectively to distinct interventional oracles, the counting bound gives Θ(n2)\Theta(n^2)5 with probability at least Θ(n2)\Theta(n^2)6; the adjacency-matrix encoding gives the matching upper bound Θ(n2)\Theta(n^2)7. A uniformly random graph thus has Θ(n2)\Theta(n^2)8 except with probability Θ(n2)\Theta(n^2)9.

A rooted-tree warm-up achieves XYX \to Y0: all trees share the same "all-equal" distribution, interventions decode descendant sets, and Cayley's formula gives XYX \to Y1 ambiguity, matched by the Prüfer encoding.

Degree-sensitive optimality

The quadratic construction uses unbounded indegree, and the paper shows this is necessary rather than incidental. For finite-gate-schema SCM classes of indegree XYX \to Y2 — uniform families built from fixed libraries of computable gate schemas (copy, AND, parity, etc.) and noise distributions — an explicit encoding (topological order, parent sets, gate/noise choices) yields

XYX \to Y3

Consequently, bounded-indegree classes admit gaps of at most XYX \to Y4, achieved by the tree construction, while dense classes admit up to XYX \to Y5, achieved by the bipartite construction. Both constructions are therefore order-optimal within this class, and the transition from XYX \to Y6 to XYX \to Y7 coincides exactly with the transition from bounded-degree to dense mechanisms.

Notably, the separation survives finite precision. Under XYX \to Y8-accurate total-variation descriptions, distinct graphs differ on some edge whose interventional marginal shifts by exactly XYX \to Y9, so the XYX \leftarrow Y0-balls around the XYX \leftarrow Y1 answer oracles are pairwise disjoint for every fixed XYX \leftarrow Y2. The quadratic lower bound persists verbatim in high probability.

Counterfactual gap beyond complete interventions

At rung 3, a modular-XOR construction stacks XYX \leftarrow Y3 independent two-variable modules, each either noise-driven (XYX \leftarrow Y4) or XOR-driven (XYX \leftarrow Y5) according to a hidden string XYX \leftarrow Y6. All modules are observationally uniform and behave identically under every hard atomic intervention — including the full hard-do oracle XYX \leftarrow Y7 over all subsets — yet counterfactuals that fix the noise distinguish the mechanisms perfectly: under no-effect, XYX \leftarrow Y8 on shared noise; under XOR they always differ. Hence even conditioning on the complete hard-do interventional oracle leaves a XYX \leftarrow Y9 counterfactual description gap. This is a strong claim: completeness of identification at rung 2 does not bound the residual rung-3 information. One caveat stated plainly: 2Θ(n2)2^{\Theta(n^2)}0 covers only hard atomic do-interventions on endogenous variables, excluding soft, stochastic, edge, and exogenous interventions, so the result does not address those richer intervention classes.

Ambiguity-to-bits and learning consequences

The constructions instantiate a general principle. Define the residual ambiguity class 2Θ(n2)2^{\Theta(n^2)}1 as the set of higher-rung answer objects consistent with lower-rung answer 2Θ(n2)2^{\Theta(n^2)}2. In Kolmogorov form, some model consistent with 2Θ(n2)2^{\Theta(n^2)}3 has gap at least 2Θ(n2)2^{\Theta(n^2)}4 minus constants, and all but a 2Θ(n2)2^{\Theta(n^2)}5 fraction exceed it; in Shannon form, if the higher-rung answer is an injective function of a hidden parameter on level sets of the lower-rung answer, the conditional entropy equals 2Θ(n2)2^{\Theta(n^2)}6 under uniform priors. Each construction's gap equals the log of its ambiguity class size (2Θ(n2)2^{\Theta(n^2)}7, 2Θ(n2)2^{\Theta(n^2)}8, 2Θ(n2)2^{\Theta(n^2)}9) matched by explicit encodings, so the gap is not a Kolmogorov-incompressibility artifact but appears equally as Shannon conditional entropy.

This yields a no-free-lunch corollary: since all MM0 bipartite mechanisms share the same observational law, any observational dataset satisfies MM1 regardless of sample size, so no learner recovers the interventional oracle with probability better than MM2, and per-query predictors incur expected absolute error at least MM3. These bounds are information-theoretic, not computational. The paper connects this to observed uneven causal reasoning in LLMs as one explanation for why predictive training alone need not induce interventional competence in worst-case structured families.

Limitations and open questions

The paper is candid about scope. All SCMs are binary and acyclic, constructed adversarially, and need not satisfy positivity or faithfulness — indeed, the constructions deliberately violate faithfulness-type assumptions, so the separations demonstrate that observational adequacy alone imposes no small causal-description bound without further structure. Whether sparsity, positivity or noise conditions, smoothness, or faithfulness shrink the residual ambiguity is left open, as are extensions beyond binary SCMs and the active-intervention complexity of closing the gap (how many experiments suffice to reduce MM4). The degree-sensitive upper bound applies only within finite-gate-schema classes with fixed libraries, leaving unrestricted SCM classes unbounded above. Finally, the counterfactual separation's restriction to hard atomic interventions means its robustness under richer intervention taxonomies remains unresolved.

Conclusion

The paper quantifies Pearl's hierarchy in bits: with constant-bit observations, the residual interventional information can be MM5, robust to constant total-variation error and order-optimal for dense gate-schema SCMs, while even the full hard-do interventional oracle can leave a MM6 counterfactual gap. The governing quantity across all three rungs is the logarithm of residual higher-rung ambiguity, established simultaneously in worst-case Kolmogorov and average-case Shannon forms.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.