---
title: Conditional Additive Independence (CAI)
url: https://www.emergentmind.com/topics/conditional-additive-independence-cai
type: topic
---

# Conditional Additive Independence (CAI)

Searching arXiv for the specified papers and closely related work on conditional additive independence.
Conditional Additive Independence (CAI) denotes a conditional independence relation defined through additivity rather than through a parametric likelihood. In the discrete graphical-model formulation of "An additive graphical model for discrete data" [2112.14674], the paper uses the term **additive conditional independence** (ACI) and states explicitly that **“conditional additive independence” (CAI) is synonymous**. The relation is defined by orthogonality of additive function spaces after conditioning, satisfies the semi-graphoid axioms, and yields a graphical model in which zeros of an inverse covariance-type operator encode conditional relations for discrete node variables without imposing an Ising or other log-linear parametric form [2112.14674]. The same acronym has also appeared in distinct literatures, including Bayesian-network analyses of noisy-or and multiattribute utility theory, where it refers to different objects; that terminological overloading is central to interpreting the phrase correctly across fields [1303.5704], [1302.1568].

## 1. Formal definition in additive function spaces

For a discrete random vector $X=(X^1,X^2,\dots,X^p)^\top$ and each node $i\in V=\{1,\dots,p\}$, let $L^2(P_{X^i})$ be the class of square-integrable mean-zero functions of $X^i$, and let $\mathscr{A}_{X^i}\subset L^2(P_{X^i})$ be a Hilbert subspace. For a nonempty subset $A\subset V$, the additive family is
\[
\mathscr{A}_{X^A}=\sum_{i\in A}\mathscr{A}_{X^i}
=\Big\{\sum_{i\in A}\phi_i:\phi_i\in \mathscr{A}_{X^i},\ i\in A\Big\}.
\]
For subspaces $\mathscr{A},\mathscr{B}\subset L^2(P_X)$, the notation $\mathscr{A}\ominus\mathscr{B}=\mathscr{A}\cap \mathscr{B}^\perp$ denotes orthogonality in $L^2(P_X)$ [2112.14674].

The formal definition given in [2112.14674] is:
\[
(\mathscr{A}_{X^A}+\mathscr{A}_{X^C})\ominus\mathscr{A}_{X^C}\ \perp\ 
(\mathscr{A}_{X^B}+\mathscr{A}_{X^C})\ominus\mathscr{A}_{X^C},
\]
and $X^A$ and $X^B$ are then said to be additively conditionally independent given $X^C$ with respect to $(\mathscr{A}_{X^A},\mathscr{A}_{X^B},\mathscr{A}_{X^C})$. This definition preserves the conditional-separation intuition of ordinary conditional independence, but the notion of “irrelevance” is formulated in terms of additive function spaces rather than factorization of a joint law [2112.14674].

A central point is that the relation is three-way and statistical, but not merely a restatement of standard conditional independence. The paper emphasizes that CAI/ACI is a semi-graphoid relation with an explicitly additive structure. This suggests that CAI is intended as a conditional-separation concept that interpolates between full nonparametric flexibility and the graph-theoretic machinery usually associated with conditional independence models.

## 2. Semi-graphoid structure and the DASG graph

The semi-graphoid framework adopted in [2112.14674] states that ACI satisfies the following axioms:

- **Symmetry**: $(A,C,B)\in\mathcal{R}\Rightarrow (B,C,A)\in\mathcal{R}$.
- **Decomposition**: $(A,C,B\cup D)\in\mathcal{R}\Rightarrow (A,C,B)\in\mathcal{R}$.
- **Weak union**: $(A,C,B\cup D)\in\mathcal{R}\Rightarrow (A,C\cup B,D)\in\mathcal{R}$.
- **Contraction**: $(A,C\cup B,D)\in\mathcal{R}$ and $(A,C,B)\in\mathcal{R}\Rightarrow (A,C,B\cup D)\in\mathcal{R}$.

These axioms support a graph-based representation analogous to conditional-independence graphs. For discrete $X\in\{0,1,\dots,m\}^p$, the paper defines the **discrete additive semi-graphoid (DASG) model** through the equivalence
\[
X^i \perp_A X^j \mid X^{-\{i,j\}} \ \Leftrightarrow\ (i,j)\notin\mathcal{E}.
\]
Thus, absence of an edge is equivalent to additive conditional independence given the remaining variables, and edges represent violations of that relation [2112.14674].

This construction is significant because it produces a nonparametric graphical model for discrete variables that does not inherit the binary-only or pairwise-potential restrictions of the Ising model. The graph is still read using the familiar language of separation and adjacency, but its semantics are operator-theoretic and additive rather than log-linear [2112.14674].

## 3. Operator characterization and finite-dimensional representations

The operator formulation is built from cross-covariance operators on the additive Hilbert spaces. For $\phi\in\mathscr{A}_{X^i}$ and $\psi\in\mathscr{A}_{X^j}$, the cross-covariance operator satisfies
\[
\langle\phi,\Sigma_{X^iX^j}\psi\rangle_{\mathscr{A}_{X^i}}
= \mathrm{cov}[\phi(X^i),\psi(X^j)].
\]
By Riesz representation and finite discreteness, $\Sigma_{X^iX^j}$ is a bounded linear operator $\mathscr{A}_{X^j}\to\mathscr{A}_{X^i}$ [2112.14674].

The **discrete additive variance operator (DAVO)** is
\[
\Sigma_{XX}:=\big\{\Sigma_{X^iX^j}\big\}_{i,j=1}^p,\qquad
\Sigma_{XX}\phi=\sum_{i,j=1}^p \Sigma_{X^iX^j}\phi_j
\]
for $\phi=\phi_1+\dots+\phi_p\in\mathscr{A}_X$. Under $\ker(\Sigma_{XX})=\{0\}$, the operator is invertible, and the **discrete additive precision operator (DAPO)** is
\[
\Theta_{XX}:=\Sigma_{XX}^{-1}.
\]
Let $\Theta_{X^iX^j}$ denote the $(i,j)$ block of $\Theta_{XX}$, viewed as a mapping $\mathscr{A}_{X^j}\to\mathscr{A}_{X^i}$ [2112.14674].

The key operator theorem is two-part:

- **Pairwise independence**: $X^i\perp X^j$ if and only if $\Sigma_{X^iX^j}=0$.
- **Additive conditional independence**: $X^i \perp_A X^j \mid X^{-\{i,j\}}$ if and only if $\Theta_{X^iX^j}=0$.

Off-diagonal blocks of $\Sigma_{XX}$ therefore encode unconditional pairwise independence, while off-diagonal blocks of $\Theta_{XX}$ encode additive conditional independence given the complement [2112.14674]. In that sense, DAPO plays the role that a precision matrix plays in Gaussian graphical models, but in an additive discrete setting.

For computation, the paper gives two finite-dimensional coordinate systems. In the **orthonormal representation**, each node $i$ with support $\{0,\dots,m\}$ has an orthonormal basis $\{u_i^{(1)},\dots,u_i^{(m)}\}$ of mean-zero functions in $\mathscr{A}_{X^i}$ and
\[
U_i(\cdot)=(u_i^{(1)}(\cdot),\dots,u_i^{(m)}(\cdot))^\top.
\]
Then
\[
[\Sigma_{X^iX^j}]_{\mathrm{o}}=\mathrm{cov}[U_i(X^i),U_j(X^j)]\in\mathbb{R}^{m\times m},
\qquad
[\Theta_{XX}]_{\mathrm{o}}=[\Sigma_{XX}]_{\mathrm{o}}^{-1}.
\]

In the **vertex representation**, for support $\{0,\dots,m\}$ one uses
\[
V(\cdot)=\big(\mathbbm{1}_{\{1\}}(\cdot),\dots,\mathbbm{1}_{\{m\}}(\cdot)\big)^\top,
\]
with
\[
[\Sigma_{X^iX^j}]_{\mathrm{v}}=\mathrm{cov}[V(X^i),V(X^j)]\in\mathbb{R}^{m\times m},
\qquad
[\Theta_{XX}]_{\mathrm{v}}=[\Sigma_{XX}]_{\mathrm{v}}^{-1}.
\]
If node $i$ has support size $m_i$, the block dimensions become $m_i\times m_j$ [2112.14674].

The paper proves the equivalences
\[
\Sigma_{X^iX^j}=0\Leftrightarrow [\Sigma_{X^iX^j}]_{\mathrm{o}}=0\Leftrightarrow [\Sigma_{X^iX^j}]_{\mathrm{v}}=0,
\]
\[
\Theta_{X^iX^j}=0\Leftrightarrow [\Theta_{X^iX^j}]_{\mathrm{o}}=0\Leftrightarrow [\Theta_{X^iX^j}]_{\mathrm{v}}=0.
\]
Accordingly, the zero-pattern of CAI can be read directly from matrix blocks in either coordinate system [2112.14674].

## 4. Penalized estimation, ADMM, and consistency

With $n$ i.i.d. samples, [2112.14674] estimates the DAPO in vertex representation by minimizing a convex D-trace loss with a blockwise group-lasso penalty:
\[
[\hat{\Theta}_{XX}]_{\mathrm{v}}
:= \underset{\Theta = \Theta^\top}{\arg\min}
\left\{
L_D(\Theta , [\hat{\Sigma}_{XX}]_{\mathrm{v}})
+\lambda_n \sum_{1\le i\ne j\le p}\| \Theta_{[i,j]}\|_\mathrm{F}
\right\},
\]
where
\[
L_D(\Theta,\Sigma)=\tfrac12 \langle\Theta^2,\Sigma\rangle_\mathrm{F}-\mathrm{tr}(\Theta).
\]
Here $\Theta\in\mathbb{R}^{mp\times mp}$, $\Theta_{[i,j]}$ is the $(i,j)$ block, and $\|\cdot\|_\mathrm{F}$ is the Frobenius norm. The group-lasso penalty induces exact block sparsity, so edge selection is implicit: an undirected edge $(i,j)\in\mathcal{E}$ is declared present if the $(i,j)$ block of $[\hat{\Theta}_{XX}]_{\mathrm{v}}$ is nonzero [2112.14674].

The convex program is solved by ADMM with a duplicate variable $\Theta_0$ and augmented Lagrangian
\[
\mathcal{L}(\Theta,\Theta_0,\Lambda)
=\tfrac12 \langle\Theta^2,[\hat{\Sigma}_{XX}]_{\mathrm{v}}\rangle_\mathrm{F}
-\mathrm{tr}(\Theta)
+\lambda_n \sum_{i\ne j}\| (\Theta_0)_{[i,j]}\|_\mathrm{F}
+\langle\Lambda, \Theta-\Theta_0\rangle_\mathrm{F}
+\tfrac{\rho}{2}\| \Theta-\Theta_0\|_\mathrm{F}^2.
\]
The iterative updates are:

- **$\Theta$-update**:
  \[
  \Theta^{(t+1)} = H([\hat{\Sigma}_{XX}]_{\mathrm{v}}+\rho I_{mp},\ I_{mp}+\rho\Theta^{(t)}_0-\Lambda^{(t)}),
  \]
  where
  \[
  H(A,B) := \underset{\Theta = \Theta^\top}{\arg\min}\Big\{\tfrac12\langle\Theta^2,A\rangle_\mathrm{F}-\langle \Theta,B\rangle_\mathrm{F}\Big\}.
  \]
  If $A = D_A \Sigma_A D_A^\top$ with eigenvalues $\sigma_1\ge\dots\ge \sigma_{mp}$, then
  \[
  H(A,B)=D_A\{(D_A^\top B D_A)\circ C\}D_A^\top,\qquad
  C_{ab}=\frac{2}{\sigma_a+\sigma_b}.
  \]

- **$\Theta_0$-update**:
  \[
  \Theta_0^{(t+1)}=S\Big(\Theta^{(t+1)}+\tfrac1\rho\Lambda^{(t)},\tfrac{\lambda_n}{\rho}\Big),
  \]
  with block soft-thresholding
  \[
  S(A,\lambda)_{[i,j]}=
  \begin{cases}
  A_{[i,j]} & i=j,\\
  \big(1-\lambda/\| A_{[i,j]}\|_\mathrm{F}\big) A_{[i,j]} & i\ne j,\ \| A_{[i,j]}\|_\mathrm{F}>\lambda,\\
  0 & i\ne j,\ \| A_{[i,j]}\|_\mathrm{F}\le \lambda.
  \end{cases}
  \]

- **Dual update**:
  \[
  \Lambda^{(t+1)}=\Lambda^{(t)}+\rho(\Theta^{(t+1)}-\Theta_0^{(t+1)}).
  \]

Initialization is
\[
\Theta^{(0)}=\Theta_0^{(0)}=[\mathrm{Diag}([\hat{\Sigma}_{XX}]_{\mathrm{v}})]^{-1},\qquad
\Lambda^{(0)}=0.
\]
The paper recommends choosing $\rho>0$, iterating until convergence, and returning $\Theta_0^{(t)}$ [2112.14674].

The high-dimensional theory is stated under sample-DAVO concentration and a blockwise irrepresentable condition. For any $\tau>2$, with probability at least $1-1/p^{\tau-2}$,
\[
\|[\hat{\Sigma}_{XX}]_{\mathrm{v}}-[\Sigma_{XX}]_{\mathrm{v}}\|_\infty
\le
\frac{3}{\sqrt{2}}
\sqrt{\frac{\log(6 m^2)+\tau\log(p)}{n}}.
\]
Under bounded $\kappa_\Gamma$, $\kappa_\Sigma$, and maximum degree $d$, and assuming the irrepresentable condition, there exists $C>0$ such that if
\[
n>\max\big(m^{-2},\gamma^{-2}(\kappa_\Sigma\kappa_\Gamma+1)^2\big)\cdot \frac{81}{2} \kappa_\Gamma^2 m^4 d^2\,[\log(6 m^2)+\tau\log(p)],
\]
and
\[
\lambda_n=\frac{9}{\sqrt{2}}\gamma^{-1} (\kappa_\Sigma\kappa_\Gamma^2+\kappa_\Gamma)m^{\frac52}d\,n^{-\frac12}\sqrt{\log(6m^2)+\tau\log(p)},
\]
then with probability at least $1-1/p^{\tau-2}$ the solution $[\hat{\Theta}_{XX}]_{\mathrm{v}}$ is unique and
\[
\|[\hat{\Theta}_{XX}]_{\mathrm{v}}-[\Theta_{XX}]_{\mathrm{v}}\|_\infty
\le
C\, m^{\frac52}d\,\sqrt{\frac{\log(6m^2)+\tau\log(p)}{n}}.
\]
For the binary case $m=1$, the paper also gives $\|\cdot\|_{1,\infty}$ and operator-norm rates [2112.14674].

The practical guidance is correspondingly explicit: $\lambda_n$ may be selected via cross-validation or information criteria; the ADMM iterations require an eigendecomposition of an $mp\times mp$ matrix; the vertex representation allows fast computation of $[\hat{\Sigma}_{XX}]_{\mathrm{v}}$ from indicator counts; and larger $\lambda_n$ yields exact zero blocks and therefore clean edge selection [2112.14674].

## 5. Relation to ordinary conditional independence and to parametric discrete models

A major contribution of [2112.14674] is to clarify when CAI coincides with ordinary conditional independence (CI) and when it does not. The analysis is developed in binary and symmetric Ising settings through **linear conditional mean (LCM)** and separator structure. For node $i$ and subset $D\subset V\setminus\{i\}$, $\mathscr{A}_{X^i}$ has linear conditional mean with respect to $\mathscr{A}_{X^D}$ if either $D=\emptyset$ or for any $\phi\in\mathscr{A}_{X^i}$,
\[
E(\phi(X^i)\mid X^D)\in \mathscr{A}_{X^D}.
\]
When $X$ is binary, this reduces to
\[
E(X^i-EX^i\mid X^D)=\xi^\top (X^D-EX^D)
\]
for some $\xi\in\mathbb{R}^{|D|}$ [2112.14674].

Under sparse-separator conditions, CI implies ACI. If $X\in\{a_0,a_1\}^p$ follows an MRF and there exists an $(i,j)$-separator $R$ such that all nodes in the component $C_i$ have LCM with respect to $X^R$, then
\[
X^i \perp X^j \mid X^{-\{i,j\}} \Rightarrow X^i \perp_A X^j \mid X^{-\{i,j\}}.
\]
For the symmetric Ising model, the same implication holds when there exists an $(i,j)$-separator $R$ with $|R|\le 2$ [2112.14674].

The paper also proves equivalence results. Under symmetric Ising, if $X^i$ has linear conditional mean with respect to $X^{-\{i,j\}}$, then
\[
X^i \perp X^j \mid X^{-\{i,j\}}
\ \Leftrightarrow\
X^i \perp_A X^j \mid X^{-\{i,j\}}.
\]
A simpler local criterion is that if $|\mathcal{N}_i\setminus\{j\}|\le 2$ or symmetrically for $j$, then the same equivalence holds. Global equivalence follows for graphs with degrees at most $2$, such as loops and threads [2112.14674].

When the neighborhoods are not sparse enough, CAI can be strictly stronger than CI. The paper gives a 5-node “almost complete” symmetric Ising counterexample in which $\Theta_{X^1X^2}$ has Hilbert–Schmidt norm $27/5516$, so the DASG graph is fully connected even though the CI graph omits the $(1,2)$ edge [2112.14674]. This is not presented as a defect of the formalism; rather, it identifies additive dependence that CI does not regard as an edge relation under the usual Markov interpretation.

To bridge that gap, [2112.14674] uses augmentation inspired by Loh and Wainwright (2012). For a nonempty separator $R$, define
\[
\mathscr{F}_D(X)
:=
\Big\{
\prod_{i\in D} u^{(n_i)}(X^i):
n_i \in\{0,1\},\
\sum_{i\in D}n_i>1,\
\sum_{i\in D}n_i\equiv 1\ (\mathrm{mod}\ 2)
\Big\},
\]
with $u^{(0)}(x)=1$ and $u^{(1)}(x)=x$. For
\[
Y^\top \overset{d}{=} (X^\top,\mathscr{F}_R(X)^\top),
\]
the paper shows
\[
Y^i \perp Y^j \mid Y^{-\{i,j\}}
\Rightarrow
Y^i \perp_A Y^j \mid Y^{-\{i,j\}},
\]
and if $R=\mathcal{N}_i\setminus\{j\}$, then equivalence holds:
\[
Y^i \perp Y^j \mid Y^{-\{i,j\}}
\ \Leftrightarrow\
Y^i \perp_A Y^j \mid Y^{-\{i,j\}}.
\]
A plausible implication is that augmentation restores agreement between additive and ordinary conditional notions by explicitly encoding interaction structure that would otherwise remain latent.

Compared with parametric discrete models, the DASG construction avoids the binary-only, pairwise-potential, and log-linear restrictions associated with the Ising family, while retaining a simple zero-block criterion for graphical recovery [2112.14674].

## 6. Alternative uses of the term “CAI”

The expression “Conditional Additive Independence” is not uniform across literatures. The following uses appear in the cited sources.

| Source | Usage of CAI or related term | Core condition |
|---|---|---|
| [2112.14674] | CAI is synonymous with additive conditional independence for discrete graphical models | $\Theta_{X^iX^j}=0 \Leftrightarrow X^i \perp_A X^j \mid X^{-\{i,j\}}$ |
| [1303.5704] | CAI is linked to additive interaction in Bayesian networks | $S_{\mathrm{add}}(e)=0$ at a chosen evidence state |
| [1302.1568] | CAI is the conditional form of additive independence in MAUT | $u(x,y,z)=f(x,z)+g(y,z)+h(z)$ |

In the Bayesian-network setting of "Conditional Inter-Causally Independent" Node Distributions, a Property of "Noisy-Or" [1303.5704], the focus is on binary parents $A,B$ and binary child $E$. There, “one common formalization of ‘additive independence’” is expressed through **additive synergy** at an evidence state $e$:
\[
S_{\mathrm{add}}(e)=P(e\mid A=1,B=1)+P(e\mid A=0,B=0)-P(e\mid A=1,B=0)-P(e\mid A=0,B=1).
\]
The paper distinguishes this forward-model notion from **CICI**—conditional inter-causal independence of the parents given one evidence state—characterized for binary matrices by a rank-one or zero-determinant condition. For CICI distributions, Proposition 5 gives
\[
S_{\mathrm{add}}(e+)=S_{\mathrm{mult}}(e+),
\]
and in noisy-or the informative evidence state $e+$ is exclusionary, with $S_{\mathrm{add}}(e+)<0$ [1303.5704]. The paper explicitly states that **CICI (posterior) $\neq$ CAI (forward additive separability)**. This guards against a common misreading: zero additive interaction in the child conditional table is not the same concept as additive conditional independence in a graphical model.

In "Conditional Utility, Utility Independence, and Utility Networks" [1302.1568], CAI belongs to multiattribute utility theory. There, conditional additive independence is the conditional version of additive independence: for attributes partitioned into $X,Y,Z$,
\[
u(x,y,z)=f(x,z)+g(y,z)+h(z)\qquad \text{for all } x,y,z.
\]
Within Shoham’s utility-distribution framework, conditional utility is
\[
u(A\mid B)=\frac{u(A\cap B)}{u(B)},
\]
and subjective utility independence is $x\perp_u y$ iff $u(x\mid y)=u(x)$ [1302.1568]. The paper connects CAI to this subjective conditional utility independence by stating that under CAI,
\[
X \perp_u Y \mid Z
\quad \text{iff} \quad
u(X\mid Y,Z)=u(X\mid Z).
\]
This is a different object again: the domain is preference representation rather than stochastic dependence.

Taken together, the sources show that “CAI” is an overloaded acronym. In discrete graphical modeling [2112.14674], it is an operator-defined conditional-separation relation; in noisy-or analysis [1303.5704], it refers to vanishing additive interaction at a fixed evidence state; and in utility theory [1302.1568], it denotes a conditional additive decomposition of a utility function.

## 7. Empirical evidence, use cases, and limitations

The empirical study in [2112.14674] compares the D-trace group-lasso DASG estimator (**DLasso**) with the APO-based estimator and a block-adapted graphical lasso (**GLasso**), with **SpIsing** included when DASG is equivalent to the MRF. The simulation settings are fixed at $n=300$, $p=200,400$, and $100$ replications; DLasso and GLasso are tuned by 5-fold CV, while APO uses generalized CV. The reported metrics are
\[
\mathrm{TPR} = \frac{|\hat{\mathcal{E}}\cap\mathcal{E}|}{|\mathcal{E}|},\qquad
\mathrm{TNR} = \frac{|\hat{\mathcal{E}}^c\cap\mathcal{E}^c|}{|\mathcal{E}^c|},\qquad
F_1=\frac{2|\hat{\mathcal{E}}\cap\mathcal{E}|}{|\mathcal{E}|+|\hat{\mathcal{E}}|}.
\]
For Ising models, DLasso has TPR and TNR comparable to SpIsing and higher TNR than GLasso/APO, with robust $F_1$ of approximately $46$–$47\%$. For the nonparametric DASG models, DLasso attains the highest TNR and $F_1$, with acceptable TPR; the paper also reports smaller standard errors than GLasso and higher AUC in ROC curves for the nonparametric settings [2112.14674].

The HIV antiretroviral therapy analysis uses binary mutations at $99$ protease residues from $n=702$ isolates, restricted to $p=62$ residues with at least $5$ mutations. Goodness-of-fit tests via KDSD strongly reject both general Ising and symmetric Ising assumptions, with average $p$-values approximately $0.0281$ and $0.0034$, and maxima $0.044$ and $0.008$. DASG is therefore used as the primary model [2112.14674]. Under bootstrapped stability with $100$ resamples and an edge-retention threshold of at least $95$ selections, DASG yields $48$ stable edges versus Ising’s $25$. The reported interactions include K20R–M36I linkage, E35–N37–R57 structural relations, M46I modulation of V82T/I84V activity, D30N–N88D association with reduced nelfinavir susceptibility, and L10I as a mutation hub; DASG-only edges include $(32,47)$ and $(12,19)$, which the Ising model misses [2112.14674].

The practical guidance in the paper is correspondingly targeted. DASG/CAI graphs are recommended for discrete data with multi-level nodes or binary data that fail Ising or log-linear assumptions, for high-dimensional regimes with likelihood misspecification, and for settings where dependence extends beyond pairwise potentials [2112.14674]. The paper also notes several limitations: support recovery depends on a block irrepresentable condition; equivalence with CI weakens for dense graphs and large separators; poor basis choices can inflate variance for non-binary nodes, though the vertex representation is more robust; and finite-sample regimes may require careful tuning because blockwise group-lasso can overshrink weak edges [2112.14674].

Open questions listed in [2112.14674] include adaptive constructions of $\mathscr{A}_{X^i}$ for ordered or ordinal data, principled augmentation strategies to bridge CI and ACI globally, and extensions to mixed discrete–continuous nodes. These questions delimit the current scope of CAI as a graphical-model concept: it is already a full semi-graphoid relation with estimation theory and empirical demonstrations, but its broader integration with mixed-type and globally interaction-aware models remains unfinished.

Source: https://www.emergentmind.com/topics/conditional-additive-independence-cai