Causal Abstraction
- Causal abstraction is the formal study of representing a system at multiple levels while ensuring that interventions commute across levels.
- It maps low-level variables to high-level abstractions via exact, uniform, and constructive transformations that preserve causal content.
- Recent advances extend the framework to soft interventions, approximate methods, data-driven learning, and even quantum and mechanistic interpretability.
Causal abstraction is the formal study of how two structural causal models, or more generally two compositional models, can represent the same system at different levels of granularity while preserving the causal content of the higher-level description. Across the literature, the central requirement is a commutation condition: whether one intervenes and then abstracts, or first maps the intervention and then evaluates the higher-level model, the resulting high-level behavior should agree. This idea appears in exact transformations, constructive abstractions, -abstractions, -abstractions, query-specific consistency conditions, and categorical natural-transformation formulations; more recent work extends it to soft interventions, approximate settings, learning from data, mechanistic interpretability, and even quantum compositional models (Beckers et al., 2018, Lorenz et al., 18 Feb 2026).
1. Core formalism and interventional commutation
In a standard SCM formulation, a low-level model and a high-level model are related by a state map and, typically, an intervention map. One influential formulation writes a low-level model and a high-level model together with partial and surjective maps
$\tau:\mathsf{Val}(\mathbf{V}_L)\longrightharpoonup \mathsf{Val}(\mathbf{V}_H),\qquad \omega:\mathcal{I}_L\longrightharpoonup \mathcal{I}_H,$
and requires the key commutation condition
This is the exact-transformation condition: “run then abstract” equals “abstract intervention then run” (Geiger et al., 15 Aug 2025). In probabilistic variants, the same requirement is expressed distributionally as push-forward equality under interventions: or, in the earlier exact-transformation language, for every low-level intervention (Beckers et al., 2018, Zhang, 2024).
Beckers and Halpern distinguish a hierarchy of increasingly restrictive notions. An exact -transformation applies to probabilistic causal models. A uniform -transformation applies to deterministic causal models and requires exactness to hold for every choice of prior on low-level contexts, thereby preventing differences from being hidden by the “right” distribution. A 0-abstraction derives the allowed intervention sets canonically from 1 itself, and a strong 2-abstraction takes the maximal intervention sets compatible with 3. A constructive abstraction is the special case where the macro variables cluster disjoint blocks of micro-variables via local surjections (Beckers et al., 2018).
Constructive abstraction is pervasive because it makes the low-to-high mapping explicit at the variable level. In one common formulation, the low-level variables are partitioned into blocks 4 and one chooses surjective component maps 5; the induced 6 is then defined blockwise, and the abstraction is constructive precisely when the commutation condition holds for the induced intervention map (Geiger et al., 15 Aug 2025). In linear SCMs, the corresponding abstraction map takes the form
7
with an exogenous map 8, and interventional consistency on all hard interventions yields the matrix identity
9
where 0 and 1 are the reduced-form operators of the low- and high-level linear SCMs (Massidda et al., 2024).
A distinct but closely related functional formulation is the 2-abstraction. Here one specifies a surjection on relevant low-level variables together with variable-wise range maps, and demands 3-consistency: preservation of all interventional queries. Recent work shows that, under bijective range maps, consistent 4-abstractions align with constructive 5-abstractions and with Cluster DAGs (Schooltink et al., 2024).
2. Structural restrictions, graphical variants, and equivalence results
Much of the theory concerns which kinds of high-level variables may be constructed from low-level ones without violating interventional consistency. In the linear case, strong intervention-based abstraction forces each abstract variable 6 to depend on a disjoint subset 7 of concrete variables, and the corresponding concrete blocks 8 are pairwise disjoint (Massidda et al., 2024). The same work shows that the abstract causal order constrains the admissible low-level DAGs: if 9 is any topological order of the high-level model, then there is a topological order of the low-level model in which all variables of 0 precede all of 1 whenever 2 (Massidda et al., 2024).
Graphical abstraction makes these restrictions visible at the level of DAGs. A Cluster DAG groups low-level variables into clusters, inducing edges between clusters whenever corresponding low-level edges exist between members of distinct clusters; Partial Cluster DAGs extend this by allowing a remainder set of dropped variables whose only effect is to create new confounding or mediated-adjacency edges (Schooltink et al., 2024). The central equivalence theorem states that, under faithfulness and bijective range maps, the following are equivalent: a bijective 3-consistent 4-abstraction, a Cluster DAG induced by the surjection from low-level to high-level variables, and a constructive 5-abstraction (Schooltink et al., 2024).
This equivalence matters because it allows transfer between graphical and functional reasoning. One can use structural algorithms on the graph to understand functional consistency, or employ functional learning methods to recover a valid abstraction graph. The data further states that Partial Cluster DAGs preserve mediated adjacencies and induced confounding, and that every bijective 6-consistent 7-abstraction—with any subset of variables dropped—yields exactly a Partial Cluster DAG (Schooltink et al., 2024). A plausible implication is that graphical abstraction is not merely a visualization aid but an alternative semantics for the same consistency condition.
The intervention sets themselves are a major source of distinction between frameworks. Exact transformations allow the modeler to choose arbitrary low- and high-level intervention sets together with 8; 9-abstraction instead derives them from the state map by demanding that a low-level intervention is meaningful only if it corresponds to a unique high-level intervention; strong $\tau:\mathsf{Val}(\mathbf{V}_L)\longrightharpoonup \mathsf{Val}(\mathbf{V}_H),\qquad \omega:\mathcal{I}_L\longrightharpoonup \mathcal{I}_H,$0-abstraction then takes all atomic interventions for which that correspondence exists (Beckers et al., 2018). This directly addresses a recurring concern in the literature: weak abstraction notions can appear to hold simply because the intervention family is too narrow.
3. Categorical and compositional formulations
Recent work recasts causal abstraction in category-theoretic language. The most concise slogan is that abstractions are monoidal natural transformations between query-functors (Lorenz et al., 18 Feb 2026). In this setting, a compositional model over a signature $\tau:\mathsf{Val}(\mathbf{V}_L)\longrightharpoonup \mathsf{Val}(\mathbf{V}_H),\qquad \omega:\mathcal{I}_L\longrightharpoonup \mathcal{I}_H,$1 in a symmetric monoidal, cd-, or Markov category $\tau:\mathsf{Val}(\mathbf{V}_L)\longrightharpoonup \mathsf{Val}(\mathbf{V}_H),\qquad \omega:\mathcal{I}_L\longrightharpoonup \mathcal{I}_H,$2 is a strong monoidal functor
$\tau:\mathsf{Val}(\mathbf{V}_L)\longrightharpoonup \mathsf{Val}(\mathbf{V}_H),\qquad \omega:\mathcal{I}_L\longrightharpoonup \mathcal{I}_H,$3
and queries are modeled by a second signature $\tau:\mathsf{Val}(\mathbf{V}_L)\longrightharpoonup \mathsf{Val}(\mathbf{V}_H),\qquad \omega:\mathcal{I}_L\longrightharpoonup \mathcal{I}_H,$4 with semantics given by a functor $\tau:\mathsf{Val}(\mathbf{V}_L)\longrightharpoonup \mathsf{Val}(\mathbf{V}_H),\qquad \omega:\mathcal{I}_L\longrightharpoonup \mathcal{I}_H,$5 (Lorenz et al., 18 Feb 2026). Causal models arise as a special case: if $\tau:\mathsf{Val}(\mathbf{V}_L)\longrightharpoonup \mathsf{Val}(\mathbf{V}_H),\qquad \omega:\mathcal{I}_L\longrightharpoonup \mathcal{I}_H,$6 is an acyclic directed graph with chosen input- and output-nodes, then the causal signature has one object per node and one generator $\tau:\mathsf{Val}(\mathbf{V}_L)\longrightharpoonup \mathsf{Val}(\mathbf{V}_H),\qquad \omega:\mathcal{I}_L\longrightharpoonup \mathcal{I}_H,$7 for each noninput node, interpreted as a channel $\tau:\mathsf{Val}(\mathbf{V}_L)\longrightharpoonup \mathsf{Val}(\mathbf{V}_H),\qquad \omega:\mathcal{I}_L\longrightharpoonup \mathcal{I}_H,$8 in a Markov category (Lorenz et al., 18 Feb 2026).
Within this framework, a type alignment $\tau:\mathsf{Val}(\mathbf{V}_L)\longrightharpoonup \mathsf{Val}(\mathbf{V}_H),\qquad \omega:\mathcal{I}_L\longrightharpoonup \mathcal{I}_H,$9 assigns each high-level type 0 to a low-level type 1 together with an epi deterministic map
2
A downward abstraction from a high-level query model to a low-level query model consists of a monoidal functor 3 and an epic natural transformation 4 such that, for every query 5,
6
An upward abstraction uses the same type maps together with a partial, surjective map 7 on low-level queries and the same consistency equations on the image of 8 (Lorenz et al., 18 Feb 2026).
This formulation unifies several earlier notions. Constructive causal abstraction becomes a downward abstraction with respect to abstract Do-queries; exact transformations are precisely upward abstractions for concrete Do-interventions; Q-9 consistency for functional SCMs and counterfactual queries is another upward-abstraction special case; interchange abstraction is a downward abstraction for interchange queries (Lorenz et al., 18 Feb 2026). A noteworthy claim in the paper is that, although causal abstractions are usually presented as upward abstractions, common cases may more fundamentally be understood as downward abstractions (Lorenz et al., 18 Feb 2026).
The categorical literature has developed two closely related but distinct unifications. One line formulates SCMs themselves as functors and abstractions as natural transformations in categories of probability spaces, with a commuting square in 0 whose endogenous component admits a right inverse under the Semantic Embedding Principle (D'Acunto et al., 1 Feb 2025). Another line treats causal models as Markov functors from free Markov categories generated by DAGs; there, a causal abstraction is a deterministic natural transformation between the corresponding Markov functors, and the usual 1-consistency and constructive 2-abstractions are recovered as special cases in 3 or 4 (Felekis et al., 6 Oct 2025). That same framework also states that if a rule of do-calculus holds on a high-level graphical abstraction of an ADMG, then it also holds on the original low-level graph (Felekis et al., 6 Oct 2025).
A further extension in the compositional program is component-level or mechanism-level abstraction. Here one does not abstract only at the level of composite queries, but also on the individual components of the model. A component-level abstraction between 5 and 6 consists of a functor on structures
7
and a natural transformation
8
such that the usual downward abstraction is recovered on queries (Lorenz et al., 18 Feb 2026). In the causal case, each high-level mechanism 9 is sent to a low-level subdiagram 0, and one requires
1
The associated characterization theorem states that this holds exactly when the partition is “extra-simple” and “full,” or equivalently when a cd-functor sends network diagrams to network diagrams together with a natural 2 (Lorenz et al., 18 Feb 2026).
4. Learning causal abstractions from data
A central shift in recent work is from hypothesis testing to learning. Earlier applications often assumed a proposed high-level model and asked whether the low-level system implemented it; newer work treats abstraction discovery itself as a statistical problem (Saulus et al., 17 Jun 2026).
For linear SCMs, the paper “Learning Causal Abstractions of Linear Structural Causal Models” provides a complete graphical- and parameter-level characterization under linear abstraction maps and introduces Abs-LiNGAM (Massidda et al., 2024). The setting uses two datasets: a large observational dataset 3 over low-level variables and a smaller paired dataset 4 over 5. The procedure first fits the linear map 6 by least squares on 7, thresholds 8 to recover relevant blocks, abstracts the large low-level dataset via 9, discovers the abstract DAG by a non-Gaussian LiNGAM-type method, derives forbidden paths from absent abstract edges, and finally discovers the concrete DAG using these constraints (Massidda et al., 2024). In simulated settings, the paper reports ROC–AUC on recovered concrete adjacencies, wall-clock time, and precision/recall of the inferred prior knowledge, and states that with only 0 paired samples Abs-LiNGAM ties DirectLiNGAM in AUC while reducing run-time by 30–50% for moderate 1 (Massidda et al., 2024).
A different learning route is the Semantic Embedding Principle. In the absence of interventional data or specified structural functions, “Causal Abstraction Learning based on the Semantic Embedding Principle” posits that the high-level observational distribution must lie on a subspace of the low-level one (D'Acunto et al., 1 Feb 2025). Formally, if 2 and 3 are the joint observational measures and 4 is the pushforward induced by the abstraction, SEP demands a right-inverse 5 such that
6
For constructive linear abstraction, one takes 7 with 8, so 9 lies on the Stiefel manifold
0
The linear-Gaussian learning problem then minimizes a KL divergence over the Stiefel manifold, and the paper proposes three Riemannian methods: LinSEPAL-ADMM, LinSEPAL-PG, and CLinSEPAL (D'Acunto et al., 1 Feb 2025). The empirical results reported include near-zero KL, low Frobenius error, and perfect 1 scores on synthetic experiments under full prior knowledge, together with successful recovery on resting-state fMRI mappings from 45 ROIs to 14 lobes or 8 functional networks (D'Acunto et al., 1 Feb 2025).
“Causal Optimal Transport of Abstractions” removes the assumption of fully specified SCMs and learns abstraction maps from observational and interventional data using a multi-marginal optimal transport objective with do-calculus constraints (Felekis et al., 2023). The abstraction error is
2
and the learned transport plans are coupled across interventions through penalties derived from truncated factorization relations along maximal chains in the intervention poset (Felekis et al., 2023). The paper states that the full COTA objective is jointly convex in all transport plans, and reports lower MMD and Wasserstein errors than non-causal OT baselines on synthetic lung-cancer models and on an Electric Battery Manufacturing dataset (Felekis et al., 2023).
The most explicitly discovery-oriented contribution in the data block is “Unsupervised Causal Abstractions Discovery,” which studies the complementary problem of learning a high-level model directly from low-level measurements (Saulus et al., 17 Jun 2026). The paper shows that observations generated by a low-rank graph induce latents that form a causal abstraction, proves identifiability under anchor assumptions, and proposes a differentiable objective over a bipartite factor-DAG parameterization with Gumbel-softmax structure variables and augmented-Lagrangian optimization (Saulus et al., 17 Jun 2026). It reports average MCC-Pearson 3 when at least one of the parent or child mechanism families is affine, high MCC-RDC even in fully nonlinear cases, and a mechanistic-interpretability case study in which three learned factors align with divisibility concepts in a small MLP trained on decimal digits (Saulus et al., 17 Jun 2026).
5. Mechanistic interpretability and computational explanation
Causal abstraction has become a core formalism in mechanistic interpretability. One influential claim is that it provides a theoretical foundation for the field by generalizing from mechanism replacement to arbitrary mechanism transformation, formalizing polysemantic neurons, the linear representation hypothesis, modular features, and graded faithfulness, and unifying activation and path patching, causal mediation analysis, causal scrubbing, causal tracing, circuit analysis, concept erasure, sparse autoencoders, differential binary masking, distributed alignment search, and steering (Geiger et al., 2023).
In this literature, the preferred empirical test is often interchange intervention accuracy. One aligns a high-level variable 4 with a low-level site 5, swaps the low-level values induced by a source input into a base input, and checks whether the resulting high-level behavior agrees with the corresponding intervention in the high-level model (Geiger et al., 2023, Puyin et al., 4 May 2026). This quantity is used as a practical proxy for approximate causal abstraction, but later work emphasizes that a single global IIA score collapses heterogeneous behavior across inputs (Puyin et al., 4 May 2026).
The survey “Causal Abstraction in Model Interpretability” summarizes several concrete case studies. In multiply quantified natural language inference, Geiger et al. used interchange interventions to test whether BERT implemented a Boolean composition tree and found a nearly perfect alignment in BERT but none in a BiLSTM baseline. Wu et al. used Boundless DAS to show that Alpaca implements a two-bit Boolean causal model on a simple arithmetic task (Zhang, 2024). The survey also notes the approximation notion
6
for 7-abstraction, attributing it to Beckers et al. (Zhang, 2024).
Later work turns this evaluation into diagnosis. “Bucketing the Good Apples” defines two inputs as interchange-consistent if both directed interchange tests agree for every high-level variable, builds an Interchangeability Graph over inputs, and extracts 8-quasi-cliques, with 9 fixed throughout (Puyin et al., 4 May 2026). The four-step recipe is to restrict to correctly answered inputs, find a candidate alignment, bucket the input space by interchangeability, and then train a classifier to characterize the partition (Puyin et al., 4 May 2026). On a toy logic task, the paper reports that recursively applying this method recovers a high-level hypothesis from scratch, including the hierarchy
00
and distinguishes that computation from a logically equivalent but different formula (Puyin et al., 4 May 2026).
A related pragmatic response to imperfect faithfulness is to combine multiple simple high-level models. In “Combining Causal Models for More Accurate Abstractions of Neural Networks,” several candidate causal models are aligned to GPT-2 small fine-tuned on arithmetic and Boolean tasks, and a greedy partition of the input space assigns different input subsets to different models subject to a faithfulness threshold 01 (Pîslar et al., 14 Mar 2025). The paper defines the strength of the combined hypothesis as
02
where 03 is the subset delegated to the trivial model (Pîslar et al., 14 Mar 2025). It reports a trade-off between the strength of an interpretability hypothesis and its faithfulness, measured by interchange intervention accuracy, and states that combined models dominate any individual model in the coverage–faithfulness trade-off on the reported tasks (Pîslar et al., 14 Mar 2025).
Beyond interpretability, causal abstraction is also used to analyze computational explanation. “How Causal Abstraction Underpins Computational Explanation” advances the slogan No Computation without Abstraction: a physical system implements a computation only if the computational model is an abstraction-under-translation of the physical causal model (Geiger et al., 15 Aug 2025). The same paper argues that representational vehicles correspond to the low-level blocks aligned with high-level variables and that information, use, and misrepresentation are guaranteed by the commutation of interventions (Geiger et al., 15 Aug 2025). It also notes a controversy: if one allows arbitrary non-linear translations, strengthened causal-abstraction conditions can become trivial, so implementation claims must further restrict 04 and 05, for instance to linear maps (Geiger et al., 15 Aug 2025).
6. Approximate, soft, robust, and extended regimes
Exact commutation is often too strong. “Approximate Causal Abstraction” defines a distance across levels
06
and calls 07 a 08–09 approximate abstraction of 10 when 11 (Beckers et al., 2019). The paper proves that 12–0 approximate abstraction is exactly exact 13-abstraction, studies composition theorems, and extends the framework to probabilistic causal models with expected-distance and tail-bound notions (Beckers et al., 2019). This suggests that approximation is not an ad hoc relaxation but a metric generalization of the exact case.
Soft interventions enlarge the intervention algebra beyond constant clamping. “Causal Abstraction with Soft Interventions” defines a soft intervention as replacement of structural equations by new functions with the same domain and codomain and no new parents (Massidda et al., 2022). A first notion, low soft abstraction, generalizes the hard-intervention restriction-set criterion via soft restriction sets, but the paper proves a non-uniqueness theorem showing that this does not determine a unique 14 in general (Massidda et al., 2022). It therefore introduces a stronger soft abstraction condition requiring compatibility for all exogenous and endogenous settings: 15 from which uniqueness of 16 follows (Massidda et al., 2022). In the constructive case, the paper gives an explicit intervention map: 17 thereby turning the abstract intervention into a computable lift of the low-level mechanism change (Massidda et al., 2022).
Robustness to environmental shift is addressed in “Distributionally Robust Causal Abstractions,” which replaces a fixed exogenous distribution with a 2-Wasserstein ambiguity set around the empirical environment (Felekis et al., 6 Oct 2025). The resulting min–max objective is
18
subject to 19 for 20 (Felekis et al., 6 Oct 2025). The paper provides Gaussian and empirical concentration results for selecting the radius, and reports that DiRoCA sacrifices some clean-data accuracy but remains far more stable under Huber-type contamination, nonlinearity misspecification, and intervention-map misspecification than non-robust baselines (Felekis et al., 6 Oct 2025).
A separate line concerns how abstraction quality should be measured. “Validating Causal Abstraction Metrics on Simulated Complex Systems” introduces a benchmark of ten complex systems and evaluates 32 candidate metrics across observational, functional, information-theoretic, and causal families (Méloux et al., 30 Jun 2026). The main result stated is that only causal metrics reliably discriminate valid from invalid abstractions, and only when incorporating faithfulness testing over unmapped variables (Méloux et al., 30 Jun 2026). The paper formalizes Causal Abstraction Error as an aggregated instance-level distance under paired interventions and noise, and reports that faithful CAE variants pass all discrimination tests across every system and can converge with as few as 30 sampled interventions (Méloux et al., 30 Jun 2026). It also emphasizes a misconception common in practice: if omitted low-level variables are treated as irrelevant without intervention-based testing, hidden causal paths can escape detection (Méloux et al., 30 Jun 2026).
The scope of causal abstraction is also expanding beyond classical SCMs. “Neural Causal Abstractions” defines abstractions by clustering variables and their domains, links query-specific 21-22 consistency to Pearl’s causal hierarchy, and uses Neural Causal Models and representational NCMs to learn abstractions at different levels of granularity, including image settings such as Colored MNIST (Xia et al., 2024). “Causal and Compositional Abstraction” generalizes from classical causal models to arbitrary compositional models and further to the category 23 of controlled quantum instruments, defining quantum analogues of opening-queries and interchange-queries and proposing downward abstractions from quantum circuits to high-level classical causal models as a formal account of classical explanations of PQCs (Lorenz et al., 18 Feb 2026).
Two application domains in the data block illustrate the breadth of the framework. In social theory, “Modeling Discrimination with Causal Abstraction” models race as a high-level abstraction of lower-level features, distinguishing constitutive relations among low-level social features from causal relations at the high level, and argues that resume-audit manipulations can be interpreted as tests of the causal effect of the high-level variable 24 on 25, conditional on an accepted alignment and causal-consistency assumption (Mossé et al., 14 Jan 2025). In networked settings, “The Causal Abstraction Network” organizes Gaussian SCMs into a sheaf-like structure whose edge restrictions are transposes of constructive linear causal abstractions, with global sections characterized by the kernel of a graph Laplacian and local learning solved by the SPECTRAL algorithm (D'Acunto et al., 25 Sep 2025). A plausible implication is that causal abstraction is increasingly functioning not only as a relation between two models, but as infrastructure for multi-model systems.