Papers
Topics
Authors
Recent
Search
2000 character limit reached

Mutual Information Decomposition

Updated 14 July 2026
  • Mutual Information Decomposition is a framework that breaks down total mutual information into interpretable components like redundancy, uniqueness, and synergy.
  • It leverages methods such as bivariate PID, antichain lattice formulations, and optimization-based approaches like BROJA to quantify complex interactions.
  • The approach enhances analyses in diverse settings—including dynamical systems, complex-valued channels, and quantum regimes—by tailoring decomposition to specific informational structures.

Searching arXiv for recent and foundational work on mutual information decomposition, especially PID and related decomposition frameworks. {"8query8 information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8", "8max_results8 8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8query8} {"8query8 information decomposition\" OR 8all:\8 information decomposition\"8 OR ti:\8"Mutual Information Decomposition\"8 OR ti:\8"Partial Information Decomposition\"", "8max_results8 8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8query8} Mutual information decomposition denotes a family of constructions that refine a mutual-information quantity into interpretable components. In the standard multivariate Shannon setting, the central objective is to decompose the information that several source variables provide about a target into atoms such as redundant, unique, and synergistic terms, because ordinary mutual information does not capture the fine-grained interactions that arise in multivariate systems (&&&8query8&&&). The same general impulse appears in other settings as well: mutual information rate can be split into time-reversible and time-irreversible parts, mutual information over complex-valued channels into amplitude, phase, and cross terms, and Rényi mutual information into entropy-like components via inequalities rather than exact identities (&&&8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8&&&). As a result, “mutual information decomposition” is not a single formalism but a research area centered on how information is structurally allocated across variables, scales, and representations.

8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8. Scope and canonical forms

A convenient starting point is the bivariate source-target decomposition of PRESERVED_PLACEHOLDER_8query8. A bivariate information decomposition consists of nonnegative functions

PRESERVED_PLACEHOLDER_8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8^

satisfying

PRESERVED_PLACEHOLDER_8max_results8^

together with

PRESERVED_PLACEHOLDER_8query8^

Here PRESERVED_PLACEHOLDER_8all:\8^ is shared information, PRESERVED_PLACEHOLDER_8 OR all:\8^ and PRESERVED_PLACEHOLDER_8 OR ti:\8^ are unique information terms, and PRESERVED_PLACEHOLDER_8 OR ti:\8^ is synergistic or complementary information (&&&8max_results8&&&).

Beyond this canonical PID form, several distinct decomposition regimes recur in the literature.

Setting Decomposition Representative source
Bivariate Shannon PID shared / unique / synergistic (&&&8max_results8&&&)
Multivariate PID antichain-lattice information atoms (&&&8all:\8&&&)
Entropy-based decomposition partial entropy decomposition via pointwise common surprisal (&&&8 OR all:\8&&&)
Game-theoretic decomposition fair-share terms IAI_A for subsets AVA\subseteq V (&&&8 OR ti:\8&&&)
Dynamical systems PRESERVED_PLACEHOLDER_8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8query8^ (&&&8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8&&&)
Complex-valued channels amplitude + phase + cross (Xie et al., 2013)
Quantum Rényi setting decomposition inequalities (McKinlay et al., 2019)

This variety matters conceptually. In some papers the decomposition target is the Shannon mutual information itself; in others it is a mutual information rate, a Rényi generalization, or a model-dependent pointwise quantity. A plausible implication is that the field is best understood not as a search for one universal decomposition, but as a family of structurally different decompositions tuned to different objects and axioms.

8max_results8. Bivariate PID: atoms, axioms, and regularity

The bivariate case is the most extensively axiomatized. Williams–Beer style formulations impose nonnegativity and the linear consistency relations above, and many works further demand symmetry, self-redundancy, monotonicity, identity-type conditions, continuity, and additivity (&&&8max_results8&&&). In this regime, once one component is specified, the remaining three are fixed by the linear system, provided the consistency condition holds.

A major line of work studies which concrete proposals satisfy which desiderata. The survey in “Continuity and Additivity Properties of Information Decompositions” examines seven prominent bivariate decompositions: PRESERVED_PLACEHOLDER_8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8, PRESERVED_PLACEHOLDER_8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8max_results8, PRESERVED_PLACEHOLDER_8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8query8, PRESERVED_PLACEHOLDER_8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8all:\8, PRESERVED_PLACEHOLDER_8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8 OR all:\8, PRESERVED_PLACEHOLDER_8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8 OR ti:\8, and PRESERVED_PLACEHOLDER_8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8 OR ti:\8. It reports that PRESERVED_PLACEHOLDER_8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv88, PRESERVED_PLACEHOLDER_8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv89, PRESERVED_PLACEHOLDER_8max_results8query8, PRESERVED_PLACEHOLDER_8max_results8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8, and PRESERVED_PLACEHOLDER_8max_results8max_results8^ are continuous, whereas PRESERVED_PLACEHOLDER_8max_results8query8^ and the common-information-based PRESERVED_PLACEHOLDER_8max_results8all:\8^ family are not continuous; with respect to additivity, PRESERVED_PLACEHOLDER_8max_results8 OR all:\8^ and PRESERVED_PLACEHOLDER_8max_results8 OR ti:\8^ are additive, while PRESERVED_PLACEHOLDER_8max_results8 OR ti:\8, PRESERVED_PLACEHOLDER_8max_results88, PRESERVED_PLACEHOLDER_8max_results89, PRESERVED_PLACEHOLDER_8query8query8, and PRESERVED_PLACEHOLDER_8query8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8^ are not additive. Among the surveyed decompositions, only BROJA is both continuous and additive (&&&8max_results8&&&).

The BROJA unique-information functional is

PRESERVED_PLACEHOLDER_8query8max_results8^

with PRESERVED_PLACEHOLDER_8query8query8^ the set of joint laws on PRESERVED_PLACEHOLDER_8query8all:\8^ sharing the same PRESERVED_PLACEHOLDER_8query8 OR all:\8- and PRESERVED_PLACEHOLDER_8query8 OR ti:\8-marginals as PRESERVED_PLACEHOLDER_8query8 OR ti:\8. This construction is optimization-based (&&&8max_results8&&&).

By contrast, “Explicit Formula for Partial Information Decomposition” proposes a closed-form bivariate formula inspired by a do-operation. It defines a random variable PRESERVED_PLACEHOLDER_8query88^ satisfying

PRESERVED_PLACEHOLDER_8query89

and then sets

PRESERVED_PLACEHOLDER_8all:\8query8^

PRESERVED_PLACEHOLDER_8all:\8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8^

PRESERVED_PLACEHOLDER_8all:\8max_results8^

The exposition states that this formula satisfies the Williams–Beer axioms and additional properties such as additivity and continuity (&&&8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8query8&&&).

A parallel no-go result limits common-information-style constructions. “Synergy, Redundancy and Common Information” shows that for independent predictor random variables, any common-information-based measure of redundancy cannot induce a nonnegative decomposition of the total mutual information, and that any reasonable measure of redundant information cannot be derived by optimization over a single random variable (&&&8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8all:\8&&&). This establishes an important negative result: even in the bivariate case, not every apparently natural auxiliary-variable construction is compatible with a nonnegative PID.

8query8. Multivariate PID and the antichain-lattice problem

The classical multivariate PID formalism indexes information atoms by antichains of source subsets. For source variables PRESERVED_PLACEHOLDER_8all:\8query8^ and target PRESERVED_PLACEHOLDER_8all:\8all:\8, the antichain set is

PRESERVED_PLACEHOLDER_8all:\8 OR all:\8^

with partial order

PRESERVED_PLACEHOLDER_8all:\8 OR ti:\8^

Each PRESERVED_PLACEHOLDER_8all:\8 OR ti:\8^ indexes an atom PRESERVED_PLACEHOLDER_8all:\88, and subsystem mutual informations are recovered by Möbius inversion over the lattice (&&&8all:\8&&&).

Earlier lattice analyses already showed that the formal representation is not neutral. In information-gain lattices, redundancy components are invariant across decompositions, whereas unique and synergy components are decomposition-dependent; information-loss lattices exchange these roles, and dual decompositions were introduced to overcome the asymmetry between invariant and decomposition-dependent components (&&&8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8 OR ti:\8&&&). Maximum-entropy constructions generalized bivariate redundancy and unique information by imposing marginal-preserving and co-information constraints, together with rooted tree-based decompositions of multivariate mutual information (&&&8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8 OR ti:\8&&&).

Recent work sharpens the negative picture. “Multivariate Partial Information Decomposition: Constructions, Inconsistencies, and Alternative Measures” reports three main contributions: explicit closed-form formulas for all two-source PID atoms that satisfy the full set of axioms and desirable properties; a three-variable counterexample in which the sum of atoms exceeds the total information; and an impossibility theorem stating that no lattice-based decomposition can be consistent for all subsets when the number of sources exceeds three (&&&8query8&&&). “Structural Impossibility of Antichain-Lattice Partial Information Decomposition” strengthens this by arguing that the obstruction is representational rather than merely axiomatic. Its central theorem states that for PRESERVED_PLACEHOLDER_8all:\89 there exists no function

PRESERVED_PLACEHOLDER_8 OR all:\8query8^

such that

PRESERVED_PLACEHOLDER_8 OR all:\8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8^

for every joint distribution of PRESERVED_PLACEHOLDER_8 OR all:\8max_results8^ (&&&8all:\8&&&).

The same paper exhibits a special target-free three-variable construction, System Information Decomposition (SID), that avoids the full antichain lattice by working on a restricted half-lattice PRESERVED_PLACEHOLDER_8 OR all:\8query8^ and replacing the usual whole-equals-sum-of-parts constraint with a modified relation containing a single subtraction term. In this setting, an operational redundancy based on Gács–Körner common information yields a self-consistent three-variable entropy decomposition (&&&8all:\8&&&). This suggests that consistency can sometimes be recovered only after changing the indexing scheme or the reconstruction rule.

8all:\8. Alternative representations: entropy, games, and functions

One major alternative to standard PID is to decompose entropy first and recover mutual-information structure second. The Partial Entropy Decomposition (PED) applies PID-style Möbius inversion to multivariate entropy using a redundancy measure based on pointwise common surprisal,

PRESERVED_PLACEHOLDER_8 OR all:\8all:\8^

where PRESERVED_PLACEHOLDER_8 OR all:\8 OR all:\8^ is the local co-information. Averaging yields an entropy-redundancy function PRESERVED_PLACEHOLDER_8 OR all:\8 OR ti:\8, and Möbius inversion over the antichain lattice gives partial entropy terms PRESERVED_PLACEHOLDER_8 OR all:\8 OR ti:\8^ (&&&8 OR all:\8&&&). In the bivariate case,

PRESERVED_PLACEHOLDER_8 OR all:\88^

so mutual information appears as redundant entropy minus synergistic entropy. The PED literature also distinguishes mechanistic redundancy, related to the function of the system, from source redundancy, arising from dependencies between inputs (&&&8 OR all:\8&&&).

A second alternative is cooperative-game theory. “Information Decomposition based on Cooperative Game Theory” defines a decomposition with exactly PRESERVED_PLACEHOLDER_8 OR all:\89 terms PRESERVED_PLACEHOLDER_8 OR ti:\8query8, one for each subset PRESERVED_PLACEHOLDER_8 OR ti:\8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8, obtained as a Faigle–Kern Shapley value under precedence constraints. Singleton terms behave as unique-information analogs; terms with PRESERVED_PLACEHOLDER_8 OR ti:\8max_results8^ are synergy terms; and there are no explicit redundancy atoms. The construction satisfies local-positivity and identity simultaneously, which the exposition contrasts with standard PID (&&&8 OR ti:\8&&&). “Values of Games for Information Decomposition” extends this picture from the hierarchical value to random-order values and sharing values on the distributive lattice of down-sets, showing that Information Attribution is one point in a broader class of value-based decompositions (&&&8max_results8all:\8&&&).

A third alternative is Functional Information Decomposition (FID). Under a complete functional specification PRESERVED_PLACEHOLDER_8 OR ti:\8query8^ with uniform inputs and structural independence of inputs, FID defines

PRESERVED_PLACEHOLDER_8 OR ti:\8all:\8^

PRESERVED_PLACEHOLDER_8 OR ti:\8 OR all:\8^

PRESERVED_PLACEHOLDER_8 OR ti:\8 OR ti:\8^

with additive decomposition

PRESERVED_PLACEHOLDER_8 OR ti:\8 OR ti:\8^

The framework further states that for any nonempty proper subset PRESERVED_PLACEHOLDER_8 OR ti:\88^ with PRESERVED_PLACEHOLDER_8 OR ti:\89, the corresponding term PRESERVED_PLACEHOLDER_8 OR ti:\8query8^ is zero, so all non-singleton proper-subset redundancy terms vanish (&&&8max_results8 OR all:\8&&&). The paper explicitly notes that this zero-redundancy result follows under input independence and that correlated inputs require a different treatment.

These alternatives share a common theme: rather than adjusting redundancy axioms inside the classical antichain PID, they alter the underlying representation—entropy atoms, Shapley allocations, or complete functions.

8 OR all:\8. Decompositions outside the standard Shannon PID setting

Mutual information decomposition is also used in settings where the decomposed object is not the ordinary multivariate PID.

For a discrete-time bivariate Markov chain PRESERVED_PLACEHOLDER_8 OR ti:\8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8, the mutual information rate between PRESERVED_PLACEHOLDER_8 OR ti:\8max_results8- and PRESERVED_PLACEHOLDER_8 OR ti:\8query8-trajectories is

PRESERVED_PLACEHOLDER_8 OR ti:\8all:\8^

and decomposes as

PRESERVED_PLACEHOLDER_8 OR ti:\8 OR all:\8^

The irreversible term is directly related to entropy production: PRESERVED_PLACEHOLDER_8 OR ti:\8 OR ti:\8^ In this setting, PRESERVED_PLACEHOLDER_8 OR ti:\8 OR ti:\8^ is associated with the information landscape and PRESERVED_PLACEHOLDER_8 OR ti:\88^ with information flux (&&&8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8&&&).

For complex-valued channels with independent input amplitude and phase, the polar decomposition writes

PRESERVED_PLACEHOLDER_8 OR ti:\89

The three terms are the amplitude term, the phase term, and the cross term, and the cross term is negligible at high signal-to-noise ratio (Xie et al., 2013). This decomposition has been used to analyze product-APSK constellations and demapper simplification (Xie et al., 2013).

In the quantum Rényi setting, exact Shannon-style identities are replaced by decomposition inequalities. If IAI_A8query8^ and IAI_A8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8^ satisfy

IAI_A8max_results8^

then the paper proves bounds of the form

IAI_A8query8^

and similarly for IAI_A8all:\8. In the special case IAI_A8 OR all:\8, one recovers the exact identity

IAI_A8 OR ti:\8^

(McKinlay et al., 2019).

These examples broaden the meaning of decomposition. The decomposed structure can be dynamical irreversibility, channel geometry, or Rényi-order dependence, not only redundancy and synergy.

8 OR ti:\8. Pointwise, differentiable, and application-driven decompositions

A distinct strand of the literature emphasizes pointwise or computationally tractable decompositions. The IAI_A8 OR ti:\8-measure introduces a differentiable local redundancy quantity

IAI_A8

where IAI_A9 is an OR-of-ANDs event built from source realizations. Averaging yields a redundancy measure AVA\subseteq V8query8, and Möbius inversion on the redundancy lattice gives pointwise atoms AVA\subseteq V8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8. Because the measure is defined for individual realizations, it is differentiable with respect to the underlying probability mass function and obeys a target chain rule (&&&8query8query8&&&).

Diffusion models provide another pointwise decomposition. In “Interpretable Diffusion via Information Decomposition,” mutual information is written exactly in terms of denoising errors, and pointwise mutual information is estimated in a nonnegative orthogonal form,

AVA\subseteq V8max_results8^

Because the squared error decomposes over coordinates, one obtains a pixel-wise decomposition

AVA\subseteq V8query8^

with each AVA\subseteq V8all:\8^ (&&&8query8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8&&&). The paper uses these quantities for compositional relation-testing, pixel-level localization of words in images, object segmentation, and prompt interventions (&&&8query8mutual information decomposition partial information decomposition additivity continuity redundancy synergy arXiv8&&&).

A different application-oriented strategy uses machine learning to allocate information “bit by bit.” The distributed information bottleneck objective

AVA\subseteq V8 OR all:\8^

places a separate bottleneck on each component AVA\subseteq V8 OR ti:\8^ and tracks how information is admitted as AVA\subseteq V8 OR ti:\8^ varies. The resulting Pareto frontier in AVA\subseteq V8 operationally orders the predictive contribution of different sources, and the paper interprets coincident, delayed, or source-specific bits as redundant, synergistic, or unique effects, respectively (&&&8query8query8&&&).

Across these pointwise and computational strands, the same structural issue reappears: decomposition is useful when it supports localization, attribution, optimization, or intervention. At the same time, the negative multivariate results for antichain-lattice PID indicate that computational convenience does not by itself resolve the representational problem (&&&8query8&&&).

In contemporary research, the field is therefore split between two complementary programs. One continues to refine axioms, formulas, and impossibility results for PID itself; the other develops alternative decompositions—entropy-based, game-theoretic, functional, dynamical, quantum, or model-based—that are tailored to the specific object being decomposed. The accumulated evidence strongly suggests that no single lattice formalism captures all multivariate informational structure, whereas well-posed decompositions remain possible once the representational target is chosen with sufficient care.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (17)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Mutual Information Decomposition.