Interventional Bayesian Causal Discovery (IBCD)
- Interventional Bayesian Causal Discovery (IBCD) is a probabilistic approach that integrates observational and interventional data to uniquely identify causal structures and resolve observational equivalence.
- IBCD employs varied formulations—from full structural causal model inference to nonparametric cut configuration methods—to address challenges like unknown intervention targets and finite-sample limitations.
- IBCD frameworks leverage active experimental design and scalable computational strategies to efficiently quantify uncertainty and guide targeted interventions in large graphical models.
Interventional Bayesian Causal Discovery (IBCD) denotes Bayesian inference of causal structure from mixtures of observational and interventional data. Depending on the formulation, the latent object may be a DAG , a full structural causal model , intervention targets and regime-specific mechanism changes, or a downstream causal query ; the unifying principle is that interventions alter likelihoods in ways that can break observational equivalence while Bayesian inference propagates uncertainty over graphs, mechanisms, and experiments (Zhou et al., 2024, Toth et al., 2022).
1. Observational equivalence and the role of interventions
A core premise of IBCD is that observational data alone generally identify only a Markov equivalence class (MEC), not a unique causal DAG. Under the Markov and faithfulness assumptions, observational conditional independences determine adjacencies and v-structures, hence the CPDAG or essential graph, while the remaining ambiguity lies in the orientation of undirected edges inside chordal chain components (Zhou et al., 2024). This limitation is also the starting point of Bayesian optimal experimental design for causal networks, where observational data typically recover only the MEC and interventions are needed to resolve the directionality of uncompelled edges (Zemplenyi et al., 2021).
The foundational Bayesian logic is older than modern DAG-space samplers. In a probability-tree formulation of causal induction, the latent causal hypothesis itself is a random variable, and intervention changes the likelihood asymmetrically across competing causal orderings (Ortega, 2011). In the worked two-variable example, observing and leaves the posterior unchanged at , whereas intervening to set yields , because the manipulated likelihood differs between and 0 (Ortega, 2011). In modern terms, this is the basic mechanism by which interventional Bayesian causal discovery escapes observational symmetry.
Recent asymptotic analysis clarifies that this is not merely a finite-sample inconvenience. In a bivariate Gaussian linear SEM with heteroscedastic noise, purely observational Bayesian model selection over the two connected directions does not consistently identify the true causal structure; the posterior limit is governed by the prior on the true model relative to the push-forward prior of the alternative under the observational equivalence map (Lungu et al., 27 Mar 2026). Once interventional samples are added, connected graphs exhibit exponentially fast posterior concentration, whereas the independence graph retains a polynomial 1 rate (Lungu et al., 27 Mar 2026). This makes interventions not only useful but asymptotically decisive for Bayesian graph discrimination.
2. What is made Bayesian in IBCD
IBCD is not tied to a single latent variable or a single posterior factorization. In Active Bayesian Causal Inference, the latent object is an SCM
2
and a causal target is represented as a query 3, yielding the induced posterior
4
This formulation treats causal discovery, partial graph discovery, full model learning, and causal reasoning as posterior inference about different functions of the same unknown SCM (Toth et al., 2022).
Other IBCD formulations deliberately avoid a full posterior over all DAGs. In the finite-intervention-sample regime studied by (Zhou et al., 2024), the latent discrete object is not directly the full DAG, but, for each intervention target set 5, the cut configuration of edges crossing 6. A prior is placed over valid cut configurations 7, interventional samples 8 provide the likelihood, and the method maintains posteriors over configurations separately for each intervention target. This produces a nonparametric Bayesian procedure for orienting unresolved edges of an essential graph without positing a linear-Gaussian or additive-noise parametric SCM (Zhou et al., 2024).
In the unknown-target, general-intervention setting, the latent object becomes the tuple
9
where 0 is the observational DAG, 1 is the unknown target set in regime 2, and 3 records possibly modified parent sets induced by that intervention (Mascaro et al., 2023). The augmented 4-DAG construction adds an intervention node 5 with arrows 6 for 7, allowing conditional independences and cross-regime invariances to be expressed within a single graphical object (Mascaro et al., 2023).
A recurring source of confusion is that “Bayesian” does not always mean full posterior inference over large DAG spaces. In the Bayes-factor intervention-optimization work of (Wang et al., 2024), the method is Bayesian in its decision language and posterior weighting over hypotheses, but the Bayes factor is effectively a plug-in likelihood ratio built from fitted models 8 and 9, not a fully Bayesian marginal likelihood over structural models and parameters (Wang et al., 2024). IBCD therefore includes both fully Bayesian posterior models and narrower Bayesian design or hypothesis-discrimination procedures.
3. Intervention models: hard, soft, general, known, and unknown
The intervention side of IBCD is heterogeneous. Many formulations assume known-target hard interventions of the form
0
with truncated factorization for the interventional likelihood (Toth et al., 2022). In the sample-efficient regime of (Zhou et al., 2024), the paper assumes effectively unlimited observational information but only finitely many i.i.d. known-target 1-interventional samples, precisely to model settings in which interventions are costly and scarce (Zhou et al., 2024).
Unknown intervention targets create a coupled inference problem because the correct interventional likelihood is itself uncertain. BaCaDI addresses this by treating the causal graph, mechanism parameters, intervention targets, and intervention effects as joint latent variables, with context labels known but target masks unknown (Hägele et al., 2022). Its main exposition assumes perfect interventions, where the target mechanism is replaced and dependence on parents is removed, but the framework is stated to also accommodate soft interventions provided the intervention likelihood remains differentiable (Hägele et al., 2022).
General interventions go further by allowing interventions to modify parent sets rather than merely delete or perturb existing mechanisms. In (Mascaro et al., 2023), a valid intervention on target set 2 with induced parent sets 3 replaces
4
by
5
subject to the validity condition that the post-intervention graph remains a DAG (Mascaro et al., 2023). A notable identifiability conclusion there is that the targets 6 are identifiable from the augmented 7-DAGs even when certain induced parent-set modifications are not (Mascaro et al., 2023).
Soft interventions with unknown targets are handled score-theoretically by the interventional BGe score. In that framework, interventions are modeled as condition-specific changes in the local Gaussian regression mechanism,
8
so that parent dependence may change without being removed (Kuipers et al., 2022). The resulting local score factorizes as a product of ordinary BGe scores over intervention-defined conditions,
9
preserving decomposability and permitting posterior DAG sampling with unknown soft interventions (Kuipers et al., 2022).
A common misconception is that “interventional” automatically means “fully identifiable.” The literature repeatedly shows otherwise: DAGs and interventions may remain identifiable only up to interventional equivalence classes, especially under unknown targets or general parent-set-changing interventions (Mascaro et al., 2023). IBCD is therefore as much about calibrated posterior ambiguity as about point identification.
4. Active intervention design and query-specific discovery
A major branch of IBCD treats interventions not as fixed inputs but as design variables. In Bayesian optimal experimental design for causal networks, the next experiment 0 is chosen to reduce posterior uncertainty over structure as rapidly as possible. The key approximation replaces expected posterior entropy under hypothetical future datasets by the entropy of an intervention-induced graph partition 1, making it possible to score interventions directly from current posterior graph samples rather than repeatedly re-running posterior inference on simulated data (Zemplenyi et al., 2021). For causal DAGs, a central choice is the MEC of the intervened graph 2, which yields a theoretically justified partition for single-node interventions (Zemplenyi et al., 2021).
ABCI generalizes this design perspective beyond graph recovery. The intervention utility is query-dependent: the learner may target the full DAG 3, a graph feature 4, the full SCM 5, or an interventional quantity such as 6 (Toth et al., 2022). The design criterion is myopic mutual information,
7
where 8 is the causal target of interest (Toth et al., 2022). This directly challenges the standard two-stage workflow in which one first learns a graph and only then answers the downstream causal question; in the Bayesian view, nuisance structure should be marginalized rather than estimated purely for its own sake (Toth et al., 2022).
Not all active Bayesian intervention methods aim at full DAG learning. The method of (Wang et al., 2024) addresses the narrower problem of discriminating 9 from null alternatives via a probability of decisive and correct evidence (PDC) objective built from Bayes-factor thresholds. The intervention rule
0
maximizes the chance that the next intervention yields threshold-crossing evidence for the true hypothesis, rather than maximizing generic information gain about a graph posterior (Wang et al., 2024). The scope is therefore edge-specific causal testing rather than general structure learning.
A complementary design question is how many interventional samples are needed once targets have been chosen. In Bayesian sample size determination for causal discovery, a CPDAG and a sequence of intervention targets are taken as input, and for each ambiguous edge 1 the method chooses the smallest 2 such that the prior-predictive probability of decisive and correct Bayes-factor evidence exceeds a desired threshold 3 (Castelletti et al., 2022). The intervention-level sample size is
4
and among graph-theoretically optimal target sequences one can select the Best size Optimal Sequence by minimizing total planned sample size (Castelletti et al., 2022). This emphasizes that minimizing the number of interventions and minimizing total interventional sample cost are distinct design problems.
5. Computational strategies and scaling regimes
The computational bottleneck in IBCD is usually the joint combinatorics of DAG structure, intervention structure, and finite-sample uncertainty. One response is to reduce the latent state space. The nonparametric method of (Zhou et al., 2024) leverages the recent result of Wienöbst et al. (2023) on uniform DAG sampling in polynomial time to efficiently enumerate cut configurations and their corresponding interventional distributions for a target set, maintaining posteriors over those configurations rather than over all DAGs in a MEC (Zhou et al., 2024). This is specifically tailored to the unusual regime of effectively infinite observational information but limited intervention samples.
A second response is approximate posterior inference over continuous relaxations of discrete objects. BaCaDI introduces continuous latent variables 5 for the graph and 6 for intervention-target masks, together with a differentiable acyclicity prior and Bernoulli edge/target probabilities parameterized through smooth sigmoids (Hägele et al., 2022). The posterior over 7 is approximated by Stein Variational Gradient Descent over particles rather than by an ELBO, using Gumbel-softmax relaxations to make posterior scores differentiable (Hägele et al., 2022). In the unknown general-intervention setting, (Mascaro et al., 2023) instead develops an MCMC scheme over DAGs, intervention targets, and induced parent sets, backed by score-equivalent priors compatible with interventional equivalence classes (Mascaro et al., 2023).
Score-based Bayesian computation remains attractive when local decomposability is available. The iBGe score preserves node-wise factorization under unknown soft interventions, so it can be inserted into standard Bayesian DAG samplers and search procedures; the implementation described in the paper uses BiDAG for iterative MAP search, posterior DAG sampling, and consensus graphs (Kuipers et al., 2022). On the design side, the posterior-partition criterion of (Zemplenyi et al., 2021) is computationally notable precisely because it avoids nested posterior inference over hypothetical future datasets, reusing current posterior graph samples across all candidate interventions (Zemplenyi et al., 2021).
Recent work has also pushed IBCD toward larger graphs by changing the likelihood object. The empirical-Bayes framework explicitly named IBCD models the estimated matrix of total causal effects,
8
rather than the full data matrix, and places a matrix normal likelihood on its estimator together with a spike-and-slab horseshoe prior on edges and empirical structural priors for Erdős–Rényi and scale-free graphs (Han et al., 2 Oct 2025). This yields posterior inclusion probabilities (PIPs) for edges and is reported to scale to 9 in simulation and to a CRISPR perturbation dataset on 521 genes (Han et al., 2 Oct 2025). The same source also notes an important methodological omission: the main text frames the problem as DAG discovery but does not clearly specify how acyclicity is enforced during posterior sampling (Han et al., 2 Oct 2025).
6. Empirical domains, recurring limitations, and frontier questions
Empirically, IBCD has been evaluated on simulated DAGs, protein-signaling networks, gene-expression perturbation data, and related multi-environment settings. The unknown-general-intervention formulation is motivated by scientific contexts such as neuroimaging effective connectivity and biological network changes across conditions (Mascaro et al., 2023). The iBGe framework targets mixed observational and interventional continuous data with soft, possibly unknown interventions and is evaluated both in simulation and on protein-expression data, including the Sachs signaling benchmark (Kuipers et al., 2022). Large-scale empirical Bayes IBCD has been applied to Perturb-seq data on 521 genes, where edge posterior inclusion probabilities are used to identify robust graph structure and a scale-free prior yields stronger reproducibility across folds and across two related CRISPR screens than an Erdős–Rényi prior (Han et al., 2 Oct 2025).
Several recurring empirical patterns cut across the literature. In the limited-sample known-target regime, maintaining posteriors over cut configurations rather than over full DAGs yields superior accuracy measured by structural Hamming distance and can be adapted to answer causal questions such as estimating the causal effect of a variable that cannot be intervened (Zhou et al., 2024). In the unknown soft-intervention Gaussian setting, Bayesian score-based learning with iBGe substantially outperforms UT-IGSP in the reported simulations and produces posterior uncertainty over both graphs and intervention effects (Kuipers et al., 2022). By contrast, latent-variable discovery from high-dimensional observations remains difficult: Decoder BCD reports that observational and/or interventional data with single-node or multi-node interventions are not sufficient to orient even a single edge in the unsupervised latent setting, whereas known intervention targets used as labels materially improve recovery (Subramanian et al., 2022).
Several controversies are best read as boundary conditions for the field rather than contradictions. First, “Bayesian” is used unevenly: some methods perform posterior inference over DAGs or latent intervention structures, whereas others are Bayesian-inspired, relying on edge-wise beliefs or plug-in Bayes factors. FedCDI, for example, is a belief-based federated framework with edge-wise Bernoulli beliefs and intervention-aware aggregation, not a coherent posterior over DAGs (Abyaneh et al., 2022). Likewise, the PDC-based method is Bayesian in its decision criterion and posterior weighting over hypotheses, but not in the sense of full posterior inference over structural models (Wang et al., 2024). Second, many core IBCD papers assume causal sufficiency, perfect interventions, or known targets, which are often unrealistic in biological perturbation studies (Zhou et al., 2024, Hägele et al., 2022). Third, adjacent work shows that post-treatment selection, latent confounding, and count-valued measurement error can invalidate standard interventional DAG formulations. Fine-grained interventional equivalence classes represented by F-PAGs have been proposed for latent confounders plus post-treatment selection (Luo et al., 30 Sep 2025), and latent linear DAG models with Poisson measurement error have been developed for interventional count data, though in a frequentist rather than Bayesian framework (Zhang et al., 26 Mar 2026).
Taken together, these strands define IBCD less as a single algorithmic family than as a research program. Its central questions are how to represent causal and interventional ambiguity probabilistically, how to exploit interventions without assuming infinite samples or perfect target knowledge, how to choose informative experiments, and how to scale posterior reasoning beyond small graphs without discarding uncertainty.