---
title: Mutual Information Decomposition
url: https://www.emergentmind.com/topics/mutual-information-decomposition
type: topic
---

# Mutual Information Decomposition

Searching arXiv for recent and foundational work on mutual information decomposition, especially PID and related decomposition frameworks.
{"query":"mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv", "max_results": 10}
{"query":"all:\"mutual information decomposition\" OR all:\"partial information decomposition\" OR ti:\"Mutual Information Decomposition\" OR ti:\"Partial Information Decomposition\"", "max_results": 10}
Mutual information decomposition denotes a family of constructions that refine a mutual-information quantity into interpretable components. In the standard multivariate Shannon setting, the central objective is to decompose the information that several source variables provide about a target into atoms such as redundant, unique, and synergistic terms, because ordinary mutual information does not capture the fine-grained interactions that arise in multivariate systems [2508.05530]. The same general impulse appears in other settings as well: mutual information rate can be split into time-reversible and time-irreversible parts, mutual information over complex-valued channels into amplitude, phase, and cross terms, and Rényi mutual information into entropy-like components via inequalities rather than exact identities [1707.01940]. As a result, “mutual information decomposition” is not a single formalism but a research area centered on how information is structurally allocated across variables, scales, and representations.

## 1. Scope and canonical forms

A convenient starting point is the bivariate source-target decomposition of \(I(S;Y,Z)\). A bivariate information decomposition consists of nonnegative functions
\[
SI(S;Y,Z),\quad UI^Y(S;Y\!\setminus\!Z),\quad UI^Z(S;Z\!\setminus\!Y),\quad CI(S;Y,Z),
\]
satisfying
\[
I(S;Y,Z)=SI+UI^Y+UI^Z+CI,
\]
together with
\[
I(S;Y)=SI+UI^Y,\qquad I(S;Z)=SI+UI^Z.
\]
Here \(SI\) is shared information, \(UI^Y\) and \(UI^Z\) are unique information terms, and \(CI\) is synergistic or complementary information [2204.10982].

Beyond this canonical PID form, several distinct decomposition regimes recur in the literature.

| Setting | Decomposition | Representative source |
|---|---|---|
| Bivariate Shannon PID | shared / unique / synergistic | [2204.10982] |
| Multivariate PID | antichain-lattice information atoms | [2604.03869] |
| Entropy-based decomposition | partial entropy decomposition via pointwise common surprisal | [1702.01591] |
| Game-theoretic decomposition | fair-share terms \(I_A\) for subsets \(A\subseteq V\) | [1910.05979] |
| Dynamical systems | \(\dot I^{rev}+\dot I^{irr}\) | [1707.01940] |
| Complex-valued channels | amplitude + phase + cross | [1304.0260] |
| Quantum Rényi setting | decomposition inequalities | [1912.06277] |

This variety matters conceptually. In some papers the decomposition target is the Shannon mutual information itself; in others it is a mutual information rate, a Rényi generalization, or a model-dependent pointwise quantity. A plausible implication is that the field is best understood not as a search for one universal decomposition, but as a family of structurally different decompositions tuned to different objects and axioms.

## 2. Bivariate PID: atoms, axioms, and regularity

The bivariate case is the most extensively axiomatized. Williams–Beer style formulations impose nonnegativity and the linear consistency relations above, and many works further demand symmetry, self-redundancy, monotonicity, identity-type conditions, continuity, and additivity [2204.10982]. In this regime, once one component is specified, the remaining three are fixed by the linear system, provided the consistency condition holds.

A major line of work studies which concrete proposals satisfy which desiderata. The survey in “Continuity and Additivity Properties of Information Decompositions” examines seven prominent bivariate decompositions: \(I_{\min}\), \(I_{\MMI}\), \(I_{\red}\), \(I_{\BROJA}\), \(I_{\dep}\), \(I_{\cap}^{\wedge}/I_{\cap}^{\GH}/I_{\cap}^{*}\), and \(I_{\IG}\). It reports that \(I_{\min}\), \(I_{\MMI}\), \(I_{\BROJA}\), \(I_{\dep}\), and \(I_{\IG}\) are continuous, whereas \(I_{\red}\) and the common-information-based \(I_{\cap}\) family are not continuous; with respect to additivity, \(I_{\BROJA}\) and \(I_{\cap}^{\wedge}/I_{\cap}^{\GH}/I_{\cap}^{*}\) are additive, while \(I_{\min}\), \(I_{\MMI}\), \(I_{\red}\), \(I_{\dep}\), and \(I_{\IG}\) are not additive. Among the surveyed decompositions, only BROJA is both continuous and additive [2204.10982].

The BROJA unique-information functional is
\[
UI_{\BROJA}(S;Y\!\setminus\!Z)=\min_{Q\in\Delta_P} I_Q(S;Y\mid Z),
\]
with \(\Delta_P\) the set of joint laws on \((S,Y,Z)\) sharing the same \((S,Y)\)- and \((S,Z)\)-marginals as \(P\). This construction is optimization-based [2204.10982].

By contrast, “Explicit Formula for Partial Information Decomposition” proposes a closed-form bivariate formula inspired by a do-operation. It defines a random variable \(X'\) satisfying
\[
P\!\bigl(X'=x,\,Z=z \mid Y=y\bigr)=P\bigl(X=x\mid Z=z\bigr)\,P\bigl(Z=z\mid Y=y\bigr),
\]
and then sets
\[
I_{\mathrm{uniq}(X;Z\mid Y)}=I(X';Z\mid Y),
\]
\[
I_\cap(X,Y;Z)=I(X;Z)-I_{\mathrm{uniq}(X;Z\mid Y)},
\]
\[
I_{\mathrm{syn}(X,Y;Z)}=I(X;Z\mid Y)-I_{\mathrm{uniq}(X;Z\mid Y)}.
\]
The exposition states that this formula satisfies the Williams–Beer axioms and additional properties such as additivity and continuity [2402.03554].

A parallel no-go result limits common-information-style constructions. “Synergy, Redundancy and Common Information” shows that for independent predictor random variables, any common-information-based measure of redundancy cannot induce a nonnegative decomposition of the total mutual information, and that any reasonable measure of redundant information cannot be derived by optimization over a single random variable [1509.03706]. This establishes an important negative result: even in the bivariate case, not every apparently natural auxiliary-variable construction is compatible with a nonnegative PID.

## 3. Multivariate PID and the antichain-lattice problem

The classical multivariate PID formalism indexes information atoms by antichains of source subsets. For source variables \(S=\{S_1,\dots,S_n\}\) and target \(T\), the antichain set is
\[
A(S)=\{\alpha\subseteq\mathcal P(S)\setminus\{\emptyset\}:\alpha\neq\emptyset,\ \forall A,A'\in\alpha,\ A\nsubseteq A'\},
\]
with partial order
\[
\beta \preceq_S \alpha \iff \forall A\in\alpha\ \exists B\in\beta: B\subseteq A.
\]
Each \(\alpha\in A(S)\) indexes an atom \(I_\alpha(T;S)\), and subsystem mutual informations are recovered by Möbius inversion over the lattice [2604.03869].

Earlier lattice analyses already showed that the formal representation is not neutral. In information-gain lattices, redundancy components are invariant across decompositions, whereas unique and synergy components are decomposition-dependent; information-loss lattices exchange these roles, and dual decompositions were introduced to overcome the asymmetry between invariant and decomposition-dependent components [1612.09522]. Maximum-entropy constructions generalized bivariate redundancy and unique information by imposing marginal-preserving and co-information constraints, together with rooted tree-based decompositions of multivariate mutual information [1708.03845].

Recent work sharpens the negative picture. “Multivariate Partial Information Decomposition: Constructions, Inconsistencies, and Alternative Measures” reports three main contributions: explicit closed-form formulas for all two-source PID atoms that satisfy the full set of axioms and desirable properties; a three-variable counterexample in which the sum of atoms exceeds the total information; and an impossibility theorem stating that no lattice-based decomposition can be consistent for all subsets when the number of sources exceeds three [2508.05530]. “Structural Impossibility of Antichain-Lattice Partial Information Decomposition” strengthens this by arguing that the obstruction is representational rather than merely axiomatic. Its central theorem states that for \(|S|\ge 3\) there exists no function
\[
f:\mathbb R^K\to\mathbb R
\]
such that
\[
I(S;T)=f\bigl(\pi_S^T(\beta):\beta\in A(S)\bigr)
\]
for every joint distribution of \((S,T)\) [2604.03869].

The same paper exhibits a special target-free three-variable construction, System Information Decomposition (SID), that avoids the full antichain lattice by working on a restricted half-lattice \(A^*(S)\) and replacing the usual whole-equals-sum-of-parts constraint with a modified relation containing a single subtraction term. In this setting, an operational redundancy based on Gács–Körner common information yields a self-consistent three-variable entropy decomposition [2604.03869]. This suggests that consistency can sometimes be recovered only after changing the indexing scheme or the reconstruction rule.

## 4. Alternative representations: entropy, games, and functions

One major alternative to standard PID is to decompose entropy first and recover mutual-information structure second. The Partial Entropy Decomposition (PED) applies PID-style Möbius inversion to multivariate entropy using a redundancy measure based on pointwise common surprisal,
\[
h_{\rm cs}(a_1,\dots,a_n)=\max\{c(a_1,\dots,a_n),0\},
\]
where \(c(a_1,\dots,a_n)\) is the local co-information. Averaging yields an entropy-redundancy function \(H_{\rm cs}\), and Möbius inversion over the antichain lattice gives partial entropy terms \(H_\partial(\alpha)\) [1702.01591]. In the bivariate case,
\[
I(X_1;X_2)=H_\partial(\{1\}\{2\})-H_\partial(\{12\}),
\]
so mutual information appears as redundant entropy minus synergistic entropy. The PED literature also distinguishes mechanistic redundancy, related to the function of the system, from source redundancy, arising from dependencies between inputs [1702.01591].

A second alternative is cooperative-game theory. “Information Decomposition based on Cooperative Game Theory” defines a decomposition with exactly \(2^n\) terms \(I_A\), one for each subset \(A\subseteq V\), obtained as a Faigle–Kern Shapley value under precedence constraints. Singleton terms behave as unique-information analogs; terms with \(|A|\ge 2\) are synergy terms; and there are no explicit redundancy atoms. The construction satisfies local-positivity and identity simultaneously, which the exposition contrasts with standard PID [1910.05979]. “Values of Games for Information Decomposition” extends this picture from the hierarchical value to random-order values and sharing values on the distributive lattice of down-sets, showing that Information Attribution is one point in a broader class of value-based decompositions [2207.05519].

A third alternative is Functional Information Decomposition (FID). Under a complete functional specification \(f:\mathcal X\to\mathcal Y\) with uniform inputs and structural independence of inputs, FID defines
\[
I_{\mathrm{tot}}=I(X_1,\dots,X_n;Y),
\]
\[
I_{\mathrm{ind}(X_i)}=I(X_i;Y),
\]
\[
I_{\mathrm{syn}(X\!\to Y)}=I(X_1,\dots,X_n;Y)-\sum_{i=1}^n I(X_i;Y),
\]
with additive decomposition
\[
I(X_1,\dots,X_n;Y)=\sum_{i=1}^n I(X_i;Y)+I_{\mathrm{syn}(X\!\to Y)}.
\]
The framework further states that for any nonempty proper subset \(S\subsetneq\{1,\dots,n\}\) with \(|S|>1\), the corresponding term \(\Delta_S\) is zero, so all non-singleton proper-subset redundancy terms vanish [2509.18522]. The paper explicitly notes that this zero-redundancy result follows under input independence and that correlated inputs require a different treatment.

These alternatives share a common theme: rather than adjusting redundancy axioms inside the classical antichain PID, they alter the underlying representation—entropy atoms, Shapley allocations, or complete functions.

## 5. Decompositions outside the standard Shannon PID setting

Mutual information decomposition is also used in settings where the decomposed object is not the ordinary multivariate PID.

For a discrete-time bivariate Markov chain \(Z=(X,S)\), the mutual information rate between \(X\)- and \(S\)-trajectories is
\[
\dot I=\sum_{z',z} T_z(z')\,q_z(z|z')\,
\ln\frac{q_z(z|z')}{q_x(x|x')\,q_s(s|s')},
\]
and decomposes as
\[
\dot I=\dot I^{rev}+\dot I^{irr}.
\]
The irreversible term is directly related to entropy production:
\[
\dot I^{irr}=\tfrac12(R_z-R_x-R_s).
\]
In this setting, \(\dot I^{rev}\) is associated with the information landscape and \(\dot I^{irr}\) with information flux [1707.01940].

For complex-valued channels with independent input amplitude and phase, the polar decomposition writes
\[
I(X;Y)=I(X_{||};Y)+I(X_{\angle};Y)+I(X_{||};X_{\angle}\mid Y).
\]
The three terms are the amplitude term, the phase term, and the cross term, and the cross term is negligible at high signal-to-noise ratio [1304.0260]. This decomposition has been used to analyze product-APSK constellations and demapper simplification [1304.0260].

In the quantum Rényi setting, exact Shannon-style identities are replaced by decomposition inequalities. If \(\alpha>0\) and \(\beta,\gamma\ge \tfrac12\) satisfy
\[
\frac{\alpha}{\alpha-1}=\frac{\beta}{\beta-1}+\frac{\gamma}{\gamma-1},
\]
then the paper proves bounds of the form
\[
I^↑_\gamma(A;B)_\rho \gtrless H_\beta(B)_\rho-H^↓_\alpha(B|A)_\rho,
\]
and similarly for \(I^↓_\gamma\). In the special case \(\alpha=\beta=\gamma=1\), one recovers the exact identity
\[
I(A:B)=H(B)-H(B|A)=H(A)-H(A|B)
\]
[1912.06277].

These examples broaden the meaning of decomposition. The decomposed structure can be dynamical irreversibility, channel geometry, or Rényi-order dependence, not only redundancy and synergy.

## 6. Pointwise, differentiable, and application-driven decompositions

A distinct strand of the literature emphasizes pointwise or computationally tractable decompositions. The \(sx\)-measure introduces a differentiable local redundancy quantity
\[
i_{\cap}^{sx}(t:\alpha)=\log_2\frac{p(t\mid W_\alpha(s)=\text{true})}{p(t)},
\]
where \(W_\alpha(s)\) is an OR-of-ANDs event built from source realizations. Averaging yields a redundancy measure \(I_{\cap}^{sx}\), and Möbius inversion on the redundancy lattice gives pointwise atoms \(\pi^{sx}(t:\beta)\). Because the measure is defined for individual realizations, it is differentiable with respect to the underlying probability mass function and obeys a target chain rule [2002.03356].

Diffusion models provide another pointwise decomposition. In “Interpretable Diffusion via Information Decomposition,” mutual information is written exactly in terms of denoising errors, and pointwise mutual information is estimated in a nonnegative orthogonal form,
\[
i^o(x;y)=\tfrac12\int \mathbb E_\epsilon\bigl\|\epsilon_\theta(\tilde x_\alpha)-\epsilon_\theta(\tilde x_\alpha\!\mid y)\bigr\|^2\,d\alpha.
\]
Because the squared error decomposes over coordinates, one obtains a pixel-wise decomposition
\[
i^o(x;y)=\sum_{j=1}^n i^o_j(x;y),
\]
with each \(i^o_j\ge 0\) [2310.07972]. The paper uses these quantities for compositional relation-testing, pixel-level localization of words in images, object segmentation, and prompt interventions [2310.07972].

A different application-oriented strategy uses machine learning to allocate information “bit by bit.” The distributed information bottleneck objective
\[
\mathcal L(\{\theta_i\})=\sum_{i=1}^n \beta_i I(X_i;Z_i)-I(\mathbf Z;Y)
\]
places a separate bottleneck on each component \(X_i\) and tracks how information is admitted as \(\beta\) varies. The resulting Pareto frontier in \(\bigl(\sum_i I(X_i;Z_i),\,I(\mathbf Z;Y)\bigr)\) operationally orders the predictive contribution of different sources, and the paper interprets coincident, delayed, or source-specific bits as redundant, synergistic, or unique effects, respectively [2307.04755].

Across these pointwise and computational strands, the same structural issue reappears: decomposition is useful when it supports localization, attribution, optimization, or intervention. At the same time, the negative multivariate results for antichain-lattice PID indicate that computational convenience does not by itself resolve the representational problem [2508.05530].

In contemporary research, the field is therefore split between two complementary programs. One continues to refine axioms, formulas, and impossibility results for PID itself; the other develops alternative decompositions—entropy-based, game-theoretic, functional, dynamical, quantum, or model-based—that are tailored to the specific object being decomposed. The accumulated evidence strongly suggests that no single lattice formalism captures all multivariate informational structure, whereas well-posed decompositions remain possible once the representational target is chosen with sufficient care.

Source: https://www.emergentmind.com/topics/mutual-information-decomposition