CaDyT: Causal Discovery in Dynamical Systems
- CaDyT is a framework for continuous-time causal discovery that leverages difference-based modeling and Gaussian process inference to handle irregular sampling.
- It utilizes the Algorithmic Markov Condition and MDL principle through a greedy search to ensure parsimonious and accurate causal structure identification.
- Empirical evaluations demonstrate that CaDyT outperforms traditional discrete-time methods on both regularly and irregularly sampled datasets by faithfully recovering underlying dynamics.
CaDyT (Causal Discovery for Dynamical Systems using Theoretical Score Analysis) is a methodological framework for identifying causal structure in dynamical systems that evolve in continuous time. CaDyT addresses two key limitations in prevailing approaches: the reliance on discrete-time approximations—which can prove inadequate for irregularly sampled data—and the tendency to ignore underlying system causality. By grounding its formalism in difference-based causal modeling and leveraging Gaussian process (GP) inference to model continuous-time trajectories, CaDyT enables more accurate and assumption-light causal discovery that is robust to both regular and irregular sampling (Tagliapietra et al., 16 Dec 2025).
1. Motivation and Context
Many real-world systems exhibit dynamics governed by continuous-time causal processes, yet available observations are often irregularly sampled and the underlying structure is unknown. Standard causal discovery methods frequently employ discrete-time frameworks such as Dynamic Bayesian Networks (DBNs) or assume regular sampling intervals. These approaches are limited by discretization-induced performance degradation and insufficient modeling of the genuine continuous-time causal structure, particularly in the regime of irregular or sparse sampling. CaDyT is formulated to address these issues by adopting a modeling paradigm that aligns with the continuous nature of real systems and accommodates weaker assumptions on the dynamics (Tagliapietra et al., 16 Dec 2025).
2. Difference-Based Causal Modeling
Conventional approaches often discretize the dynamics, effectively imposing strong Markovian and time-homogeneity assumptions. In contrast, CaDyT uses difference-based causal models that permit more flexible and minimal assumptions on temporal evolution. This paradigm focuses on learning the relationships between infinitesimal or finite differences of the variables, consistent with the mathematical structure of continuous-time dynamical systems. By not imposing overly restrictive state-transition constraints, difference-based causal models facilitate more faithful inference of causal structure from real system behavior, especially when observations are unevenly distributed in time (Tagliapietra et al., 16 Dec 2025).
3. Gaussian Process Inference for Continuous-Time Dynamics
To model the underlying continuous-time dynamics, CaDyT exploits exact Gaussian Process inference. GP regression offers a nonparametric Bayesian framework for interpolating and extrapolating system trajectories, allowing for natural incorporation of uncertainty quantification and principled regularization. Unlike discretized time-series models, GPs can model data observed at arbitrary—potentially irregular—time points without parameter inflation or loss of information. The continuous-time GP model aligns the statistical representation used in discovery with the intrinsic smoothness and temporal variability present in many physical, biological, or socioeconomic systems (Tagliapietra et al., 16 Dec 2025).
4. Causal Structure Search: Algorithmic Markov Condition and MDL Principle
CaDyT identifies the system’s causal structure through a greedy search. This procedure is guided by the Algorithmic Markov Condition (AMC), which formalizes conditional independence in causal inference, and the Minimum Description Length (MDL) principle, which provides a consistent framework for model comparison and penalization of complexity. The greedy search incrementally builds the candidate structure, evaluating at each step the AMC/MDL-guided scoring criterion. Models that best balance fit to the observed dynamics (GP posterior likelihood) and parsimony (MDL penalty) are retained. This approach promotes identification of causal relations that are both theoretically justified and practically effective (Tagliapietra et al., 16 Dec 2025).
5. Empirical Evaluation
The performance of CaDyT has been empirically assessed against state-of-the-art causal discovery methods on benchmark datasets comprising both regularly and irregularly sampled time series. Experimental results demonstrate that CaDyT consistently discovers causal networks that are closer to the true system dynamics in comparison to contemporary methods. Notably, the use of exact Gaussian Process inference and difference-based modeling confers robustness to challenges posed by missing or irregular data, a setting where discrete-time methods underperform (Tagliapietra et al., 16 Dec 2025). The advantage is evident in accuracy metrics that reflect both structural fidelity and dynamic prediction.
6. Distinction from Existing Methods
CaDyT diverges from dominant causal structure learning methods in several respects. Unlike DBN-based approaches, it does not require discretization of time, nor does it depend on regular spacing of observations. Its reliance on difference-based causal modeling allows for looser—and thus more realistic—assumptions regarding inter-variable dependencies across time. The integration of Algorithmic Markov Condition for independence and MDL for complexity penalization offers a systematic foundation that contrasts with ad hoc or overly parameterized discrete-time models, ensuring generalizability to novel dynamical regimes (Tagliapietra et al., 16 Dec 2025).
7. Implications and Applications
The CaDyT methodology can be applied across domains where continuous-time dynamical systems and causal discovery intersect, such as systems biology, econometrics, and engineered process control. Its performance on both regularly and irregularly sampled datasets broadens its utility to empirical regimes that were previously inaccessible to DBN-based or purely static discovery frameworks. This suggests a plausible extension to real-world applications such as analyzing biological regulatory networks, financial markets with missing data, or sensor-driven engineering processes with stochastic measurement times (Tagliapietra et al., 16 Dec 2025).
In summary, CaDyT constitutes a principled, empirically validated advance in continuous-time causal discovery, founded on difference-based models, Gaussian process inference, and algorithmically justified structure identification. It improves upon the state of the art, particularly in settings where sampling irregularity and unknown dynamics impose severe challenges on standard approaches.