---
title: Determinant-Based Mutual Information
url: https://www.emergentmind.com/topics/determinant-based-mutual-information
type: topic
---

# Determinant-Based Mutual Information

Searching arXiv for the cited papers and closely related determinant-based mutual information work.
arxiv_search query: "2205.00794 determinant mutual information log determinant entropy blind source separation"
arxiv_search results requested for:
- 2205.00794
- 1407.7165
- 2301.08164
- 2111.00496
- 2602.12346
- 2606.22301
Determinant-based mutual information denotes a family of dependence functionals in which information is represented through a log-determinant, a determinant ratio, or an operator determinant built from second-order objects such as covariance matrices, conditional error covariances, Gram matrices, or autocorrelation operators. In the literature summarized here, these constructions serve several distinct roles: an alternative mutual-information-like quantity based on covariance log-determinants for blind source separation, a sharp lower bound on Shannon mutual information based on the determinant of conditional-mean prediction error, exact Gaussian-process and linear-Gaussian conditional mutual informations written as log-determinant differences, matrix-based entropy contrasts that recover determinant forms in the \(\alpha\to1\) limit, and Fredholm-determinant expressions for continuous random fields [2205.00794, 1407.7165, 2301.08164, 2602.12346, 2606.22301, 2111.00496].

## 1. Covariance log-determinants as entropy and mutual information

A central formulation introduces a log-determinant entropy for a random vector \(x\in\mathbb R^r\) with covariance \(R_x\succ0\). For a small regularization \(\epsilon>0\),
\[
H_{LD}^{(\epsilon)}(x)
=
\frac12\,\log\det\bigl(R_x+\epsilon I_r\bigr)
+\frac r2\log(2\pi e).
\]
In the limit \(\epsilon\to0\), this coincides with the Gaussian differential entropy \(\frac12\log\det R_x+\frac r2\log(2\pi e)\), which upper-bounds the true Shannon entropy of any \(x\) with covariance \(R_x\) [2205.00794].

For jointly distributed vectors \(x\in\mathbb R^r\) and \(y\in\mathbb R^q\), the same framework defines an LD-conditional entropy through the linear-MMSE error,
\[
H_{LD}^{(\epsilon)}\bigl(y\mid_L x\bigr)
=
\frac12\,\log\det\Bigl(R_y-R_{xy}^{T}(R_x+\epsilon I_r)^{-1}R_{xy}+\epsilon I_q\Bigr)
+\frac q2\log(2\pi e),
\]
and the associated LD-mutual information
\[
I_{LD}^{(\epsilon)}(x;y)
=
H_{LD}^{(\epsilon)}(y)-H_{LD}^{(\epsilon)}(y\mid_L x)
=
H_{LD}^{(\epsilon)}(x)-H_{LD}^{(\epsilon)}(x\mid_L y).
\]
Equivalently,
\[
I_{LD}^{(\epsilon)}(x;y)
=
H_{LD}^{(\epsilon)}(x)+H_{LD}^{(\epsilon)}(y)-H_{LD}^{(\epsilon)}([x;y]).
\]

This formulation is explicitly second-order. The corresponding LD-mutual information between two vectors reflects a level of their correlation, rather than arbitrary higher-order statistical dependence. That distinction is operational: the same paper contrasts the LD-infomax criterion with ICA infomax and states that the proposed information maximization approach can separate both dependent and independent sources [2205.00794].

A different determinant construction appears in regression-based dependence analysis. Let
\[
e(Y\mid X)=Y-E[Y\mid X],
\qquad
M=E\bigl[e(Y\mid X)e(Y\mid X)^T\bigr].
\]
The scalar dependence measure is
\[
\nu=\det M,
\]
or, in normalized form,
\[
\nu(Y\mid X)=\frac{\det M}{\det\operatorname{Var}(Y)}.
\]
The resulting inequality
\[
I(X;Y)\ge \log(\nu^{-1/2})
\]
is a sharp lower bound on Shannon mutual information, with equality in the jointly Gaussian case [1407.7165].

A plausible summary is that determinant-based mutual information is not a single invariant quantity. In the cited work it includes exact Shannon-MI formulas for Gaussian models, covariance-based surrogates, and lower bounds whose determinant structure is the organizing principle rather than an assertion of universal equivalence.

## 2. Structural identities, exactness regimes, and common misunderstandings

Several algebraic properties recur across determinant formulations. In the LD framework, \(I_{LD}^{(\epsilon)}(x;y)\ge0\) for nondegenerate \(x,y\), with equality if and only if \(R_{xy}=0\). It is symmetric by construction, and if \((x,y)\) is jointly Gaussian then \(H_{LD}\) equals the Gaussian differential entropy exactly, so \(I_{LD}\) coincides with Shannon mutual information:
\[
\frac12\log\frac{\det R_x\,\det R_y}{\det R_{[x;y]}}.
\]
Thus Gaussian exactness is an explicit property, not a generic one [2205.00794].

The regression-based bound has an analogous exactness regime. Under the assumptions that \(m(x)=E[Y\mid X=x]\) is one-to-one and continuously differentiable, and that the marginal law of \(X\) is chosen so that \(M:=E[Y\mid X]\) is Gaussian, one obtains
\[
I(X;Y)=I(M;Y)
\]
and then
\[
I(X;Y)\ge
\log\Bigl\{\det\operatorname{Var}(Y)^{1/2}/\det M^{1/2}\Bigr\}
=
\log\bigl(\nu(Y\mid X)^{-1/2}\bigr).
\]
When \((X,Y)\) is jointly Gaussian, the inequality becomes an equality [1407.7165].

The same determinant structure reappears in exact conditional mutual information for linear Gaussian models. For disjoint index sets \(A,B,C\), define
\[
\Sigma_{A\mid C}
=
\Sigma_{A,A}-\Sigma_{A,C}\Sigma_{C,C}^{-1}\Sigma_{C,A},
\]
and
\[
\Sigma_{A\mid B,C}
=
\Sigma_{A,A}
-
\begin{pmatrix}\Sigma_{A,B}&\Sigma_{A,C}\end{pmatrix}
\begin{pmatrix}\Sigma_{B,B}&\Sigma_{B,C}\\\Sigma_{C,B}&\Sigma_{C,C}\end{pmatrix}^{-1}
\begin{pmatrix}\Sigma_{B,A}\\\Sigma_{C,A}\end{pmatrix}.
\]
Then
\[
I(V_A;V_B\mid V_C)
=
\log\det\bigl(\Sigma_{A\mid C}\bigr)
-
\log\det\bigl(\Sigma_{A\mid B,C}\bigr).
\]
This is an exact closed form for jointly Gaussian vectors, obtained directly from Gaussian conditional entropies [2606.22301].

A common misunderstanding is to treat every determinant expression as an exact Shannon mutual information. The cited literature does not support that identification. Exact equality is stated for jointly Gaussian variables, Gaussian processes, linear Gaussian DAGs, and Gaussian random fields under the specified constructions; the regression-error determinant supplies a lower bound; and DiME is described as a lower bound on matrix-based mutual information and as a mutual-information-like quantity rather than as Shannon mutual information itself [1407.7165, 2301.08164, 2602.12346, 2606.22301, 2111.00496].

## 3. LD-infomax and determinant maximization in blind source separation

In noiseless blind source separation, one observes mixtures
\[
y(k)=H_g\,s_g(k)
\]
and seeks estimates \(S=[\,s(1)\;\cdots\;s(N)\,]\) constrained to lie in a known polytope \(\mathcal P\). The LD-infomax criterion is
\[
\max_{S\in\mathbb R^{r\times N}}
\;\hat I_{LD}^{(\epsilon)}(Y;S)
\quad\text{s.t.}\quad
S_{:,j}\in\mathcal P\quad(j=1,\dots,N),
\]
where the deterministic objective is formed from sample covariances
\[
\hat R_s=\tfrac1N\,S\,S^T-\tfrac1{N^2}S\mathbf1\mathbf1^TS^T,
\]
\[
\hat R_y=\tfrac1NYY^T-\tfrac1{N^2}Y\mathbf1\mathbf1^TY^T,
\]
\[
\hat R_{sy}=\tfrac1N\,S\,Y^T-\tfrac1{N^2}S\mathbf1\mathbf1^T Y^T,
\]
yielding
\[
\hat I_{LD}^{(\epsilon)}(Y;S)
=\frac12\log\det\bigl(\hat R_s+\epsilon I_r\bigr)
-\frac12\log\det\!\bigl(\hat R_s
-\hat R_{sy}(\hat R_y+\epsilon I_M)^{-1}\hat R_{sy}^T
+\epsilon I_r\bigr).
\]
This gives an information-theoretic perspective for determinant maximization-based structured matrix factorization methods such as nonnegative and polytopic matrix factorization [2205.00794].

In the limit \(\epsilon\to0\), the noiseless case reduces to
\[
\max_{H,S}\;\log\det\!\bigl(\tfrac1N\,S\,S^T\bigr)
\quad\text{s.t.}\quad
Y=H\,S,\;S_{:,j}\in\mathcal P,
\]
which is the usual determinant-maximization PMF criterion. The reduction is exact in the sense stated in the source summary: the second term is driven to \(-\infty\) unless \(Y\) and \(S\) are exactly linearly related [2205.00794].

The finite-sample perfect-separation guarantee invokes the Polytopic Matrix Factorization identifiability result. If \(\mathcal P\) is an identifiable polytope and the true source columns \(S_g\subset\mathcal P\) are sufficiently scattered in the sense that
\[
\mathrm{conv}(S_g)\supset\mathcal E_{\mathcal P},
\qquad
\mathrm{conv}(S_g)^{*\!,g_{\mathcal P}}
\cap\mathrm{bd}(\mathcal E_{\mathcal P}^{*\!,g_{\mathcal P}})
=
\mathrm{ext}(\mathcal P^{*\!,g_{\mathcal P}}),
\]
then any optimizer \(S_*\) satisfies
\[
S_*=D\,\Pi\,S_g,
\]
where \(D\) is a diagonal sign matrix and \(\Pi\) a permutation. Perfect recovery is therefore guaranteed up to the unavoidable sign-permutation ambiguity from a finite sample \(N\) that yields a sufficiently scattered configuration [2205.00794].

## 4. Regression-error determinants as sharp lower bounds

The regression-based formulation of Bowsher and Voliotis defines dependence through the determinant of the second-moment matrix of the conditional mean prediction error. With
\[
e(Y\mid X)=Y-E[Y\mid X],
\qquad
M=E[e(Y\mid X)e(Y\mid X)^T],
\]
the determinant \(\nu=\det M\) quantifies the residual uncertainty in \(Y\) after conditioning on \(X\). The lower bound
\[
I(X;Y)\ge \log(\nu^{-1/2})
\]
is derived by introducing \(M:=E[Y\mid X]\), assuming \(M\) is an invertible, continuously differentiable transform of \(X\) with Gaussian marginal law, and then applying a covariance-determinant bound to \(I(M;Y)\) [1407.7165].

The derivation uses the identity
\[
\operatorname{Cov}(Y,M)=\operatorname{Var}(M),
\]
which implies
\[
\operatorname{Var}(Y)-\operatorname{Cov}(Y,M)\operatorname{Var}(M)^{-1}\operatorname{Cov}(Y,M)^T
=
E[\operatorname{Var}(Y\mid X)].
\]
Consequently, the determinant in the bound is the determinant of the usual law-of-total-variance residual covariance. The paper further states that the bound is tighter than lower bounds based on the Pearson correlation and ones derived using average mean square-error rate distortion arguments [1407.7165].

The comparison is explicit in the bivariate case:
\[
I(X;Y)\ge \log\bigl[(1-\operatorname{Corr}^2(X,Y))^{-1/2}\bigr].
\]
Because \(\det M\le \operatorname{Var}(Y)\cdot(1-\operatorname{Corr}^2)\), the determinant-based bound is at least as large, and strictly larger whenever \(\operatorname{Corr}^2\) fails to capture non-linear or higher-order dependence. The same summary states that the determinant-based bound is strictly tighter than the average-MSE bound by an AM-GM inequality on eigenvalues [1407.7165].

Estimation proceeds by nonparametric regression: estimate \(m(x)=E[Y\mid X=x]\), form the empirical second-moment matrix of residuals,
\[
\hat M=(1/n)\sum_i [Y_i-\hat m_i][Y_i-\hat m_i]^T,
\]
and compute \(\hat\nu=\det\hat M\), with plug-in lower bound
\[
\widehat L=\log(\hat\nu^{-1/2}).
\]
The method includes BC\(_a\) bootstrap intervals for \(L\) and a composite estimator
\[
\hat I_{\mathrm{comp}}
=
\max\{\hat I_{knn},\;\text{lower bound of }(1-\alpha)\times100\%\text{ BC}_a\text{ interval for }L\},
\]
which substantially improves upon inference about mutual information based on \(k\)-nearest neighbour estimators alone in the simulations described in the source summary [1407.7165].

## 5. Matrix-based entropies, Gram determinants, and DiME

A distinct line of work defines matrix-based Rényi entropies from normalized Gram matrices. Given samples \(X=\{x_i\}_{i=1}^n\) and a positive-definite kernel \(\kappa\) with \(\kappa(x,x)=1\), form the Gram matrix \(K_X\) and its eigenvalues \(\lambda_1,\dots,\lambda_n\). The \(\alpha\)-order matrix-based Rényi entropy is
\[
S_\alpha(K_X)
=
\frac{1}{1-\alpha}\log\Bigl[\sum_{i=1}^n \lambda_i^\alpha\Bigr]
=
\frac{1}{1-\alpha}\log\operatorname{Tr}[K_X^\alpha].
\]
For the trace-normalized matrix \(\tilde K=K_X/\operatorname{Tr}(K_X)\), one has \(\sum_i\lambda_i=1\). In the limit \(\alpha\to1\),
\[
S_1(\tilde K)=-\sum_i\lambda_i\log\lambda_i,
\]
and, up to additive constants,
\[
S_1(K_X)\simeq -\log\det(K_X+\epsilon I)
\]
for small \(\epsilon\), recovering the familiar log-determinant formula from Gaussian information [2301.08164].

For two views with kernels \(K_X\) and \(K_Y\), the joint Gram matrix is defined by the Hadamard product
\[
K_{XY}=K_X\circ K_Y,
\]
and the matrix-based mutual information of order \(\alpha\) is
\[
I_\alpha(K_X;K_Y)
=
S_\alpha(K_X)+S_\alpha(K_Y)-S_\alpha(K_X\circ K_Y).
\]
In the \(\alpha\to1\) limit this recovers a log-determinant form analogous to
\[
I(X;Y)=\tfrac12\log\frac{\det\Sigma_X\det\Sigma_Y}{\det\Sigma_{XY}}.
\]

DiME is defined by contrasting the joint entropy of paired samples with that of negatively paired samples obtained by permuting one view:
\[
S_\alpha^{+}=S_\alpha(K_X\circ K_Y),
\qquad
S_\alpha^{-}=E_{\Pi}[S_\alpha(K_X\circ(\Pi K_Y\Pi^T))],
\]
\[
I_{DiME}(X;Y)=S_\alpha^{-}-S_\alpha^{+}.
\]
The same source states that this simpler form shows that DiME is a lower bound on \(I_\alpha(K_X;K_Y)\). It also gives a collapse-avoidance property: if one view collapses so that \(K_Y=\mathbf 1\mathbf 1^T\), then \(S_\alpha^{+}=S_\alpha^{-}\) and \(I_{DiME}=0\), so any collapse yields zero objective [2301.08164].

Finite-sample computation is spectral: choose kernels, compute Gram matrices, normalize, form \(K_{XY}\), diagonalize \(K_{XY}\) and permuted copies, evaluate \(S_\alpha\), and average over a small set of random permutations. A single eigendecomposition costs \(O(n^3)\), low-rank approximations such as Nyström or random features reduce this to \(O(nm^2)\) with \(m\ll n\), and the gradient is obtained through
\[
\frac{\partial S_\alpha(K)}{\partial K}
=
\frac{\alpha}{1-\alpha}\frac{K^{\alpha-1}}{\operatorname{Tr}[K^\alpha]}
\]
followed by the chain rule. The paper reports use cases in multiview representation learning and latent factor disentanglement, and compares DiME against InfoNCE, NWJ, JS, CLUB, MINE, CKA, HSIC, TUBA, DoE, and plain matrix-based MI in the specific experiments summarized in the source text [2301.08164].

## 6. Gaussian-process and linear-Gaussian network formulations

In Gaussian-process information gathering, determinant-based mutual information is exact. For a finite candidate set \(\mathcal V\) with prior covariance \(\mathbf K_{\mathcal V\mathcal V}\), and a selected subset \(\mathcal A\subset\mathcal V\), the MI between noisy observations at \(\mathcal A\) and the rest of the field is
\[
\mathbb I(\mathcal A;\mathcal V\setminus\mathcal A)
=
\tfrac12\Bigl[
\ln|\mathbf K_{\mathcal A\mathcal A}+\sigma_n^2\mathbf I_s|
+
\ln|\mathbf K_{(\mathcal V\setminus\mathcal A)(\mathcal V\setminus\mathcal A)}+\sigma_n^2\mathbf I_{m-s}|
-
\ln|\mathbf K_{\mathcal V\mathcal V}+\sigma_n^2\mathbf I_m|
\Bigr].
\]
Equivalently,
\[
\mathbb I(\mathcal A;\mathcal V\setminus\mathcal A)
=
\tfrac12\ln
\frac{\det\mathbf K_{\mathcal A\mathcal A}\,
\det\mathbf K_{(\mathcal V\setminus\mathcal A)(\mathcal V\setminus\mathcal A)}}
{\det\mathbf K_{\mathcal V\mathcal V}}.
\]
The standard exact evaluation costs \(\mathcal O(m^3)\) per determinant of size \(m\) [2602.12346].

Schur-MI exploits the iterative structure of robotic information gathering and the Schur-complement determinant identity
\[
\left|
\begin{pmatrix}
A & B\\
B^T & C
\end{pmatrix}
\right|
=
|C|\;|A-BC^{-1}B^T|.
\]
With the conditional covariance
\[
\mathbf K_{\mathcal A\mid\mathcal V}
=
\mathbf K_{\mathcal A\mathcal A}
-
\mathbf K_{\mathcal A\mathcal V}\mathbf K_{\mathcal V\mathcal V}^{-1}\mathbf K_{\mathcal V\mathcal A},
\]
one obtains
\[
\mathbb I(\mathcal A;\mathcal V)
=
\tfrac12\ln
\frac{\det\mathbf K_{\mathcal A\mathcal A}}
{\det\mathbf K_{\mathcal A\mid\mathcal V}}.
\]
After precomputing \(\mathbf K_{\mathcal V\mathcal V}^{-1}\), each evaluation requires only two \(s\times s\) determinants, reducing the per-evaluation cost from \(\mathcal O(|\mathcal V|^3)\) to \(\mathcal O(|\mathcal A|^3)\). The paper states that MI is submodular, so the greedy algorithm achieves a \((1-1/e)\) approximation to the optimal NP-hard sensor placement problem, and reports up to a \(12.7\times\) speedup over the standard log-det formulation while producing identical MI values [2602.12346].

For multi-terminal linear Gaussian wireless networks, conditional mutual information is also a log-determinant difference of block Schur complements:
\[
I(V_A;V_B\mid V_C)
=
\log\det(\Sigma_{A\mid C})-\log\det(\Sigma_{A\mid B,C}).
\]
The node-pair covariances are produced by a topologically ordered K-recursion,
\[
K_{jk}=\sum_{i\in\mathrm{Pa}(j)}A_{ji}K_{ik},
\qquad
K_{jj}=\sum_{i,i'\in\mathrm{Pa}(j)}A_{ji}K_{ii'}A_{ji'}^H+\Sigma_j,
\]
after which all computations are expressed through automatic-differentiation primitives: matrix products, Hermitian transposes, Cholesky decompositions, triangular solves, and log-determinants. A single reverse-mode sweep yields the Wirtinger gradient with respect to all controllable factors at once, allowing projected gradient iterations for weighted sum-rate, secrecy, and non-linear composites built from finitely many conditional MIs [2606.22301].

## 7. Random fields, Fredholm determinants, and spatial spectra

For continuous electromagnetic fields, determinant-based mutual information is formulated at the operator level. A current density \({\bf J}({\bf s})\) on a source region \(V_s\) generates an electric field
\[
{\bf E}({\bf r})=\int_{V_s}\mathbf G({\bf r},{\bf s})\,{\bf J}({\bf s})\,d{\bf s},
\qquad {\bf r}\in V_r,
\]
and the receiver observes
\[
{\bf Y}({\bf r})={\bf E}({\bf r})+{\bf N}({\bf r}),
\]
with \({\bf J}\) and \({\bf N}\) modeled as zero-mean Gaussian random fields. Their second moments define signal and noise autocorrelation operators \(C_s\) and \(C_n\) on \(\mathscr L^2(V_r)\) [2111.00496].

When the field autocorrelation kernel is continuous or Hilbert-Schmidt, Mercer’s theorem yields eigenfunctions \(\{\phi_k\}\) and eigenvalues \(\{\lambda_k\}\) such that
\[
R_E(r,r')
=
\sum_{k=1}^{\infty}\lambda_k\,\phi_k(r)\phi_k(r')^*,
\]
and
\[
E(r)=\sum_{k=1}^{\infty}\xi_k\,\phi_k(r),
\qquad
E[\xi_k\xi_\ell^*]=\lambda_k\delta_{k\ell}.
\]
Under spatially white Gaussian noise \(C_n=\sigma^2 I\), the mutual information per channel use becomes
\[
I=\log\det(I+C_sC_n^{-1})
=
\sum_{k=1}^{\infty}\log\Bigl(1+\frac{\lambda_k}{\sigma^2}\Bigr),
\]
equivalently
\[
\det(I+C_sC_n^{-1})=\prod_k\Bigl(1+\frac{\lambda_k}{\sigma^2}\Bigr).
\]
This is a Fredholm determinant expression for a continuous-field Gaussian channel [2111.00496].

The same source gives two extensions. First, for rational-spectrum kernels one may define
\[
f(z)=\det(I+zC_s)=\prod_{k=1}^{\infty}(1+z\lambda_k),
\]
and derive analytic closed forms without listing the individual eigenvalues. For the one-dimensional exponential kernel
\[
R_E(r,r')=P\,e^{-\alpha|r-r'|},\quad r,r'\in[0,L],
\]
the mutual information is
\[
I
=
\log\Biggl[
\cosh\!\Bigl(\alpha L\sqrt{1+\frac{4P}{\alpha\sigma^2}}\Bigr)
+
\frac{1+2P/(\alpha\sigma^2)}{\sqrt{1+4P/(\alpha\sigma^2)}}
\sinh\!\Bigl(\alpha L\sqrt{1+\frac{4P}{\alpha\sigma^2}}\Bigr)
\Biggr]
-\alpha L.
\]
Second, for colored noise one forms
\[
C_n^{-1/2}C_sC_n^{-1/2},
\]
whose generalized eigenvalues \(\{\mu_k\}\) yield
\[
I
=
\log\det\bigl(I+C_n^{-1/2}C_sC_n^{-1/2}\bigr)
=
\sum_{k=1}^{\infty}\log(1+\mu_k).
\]
In the stationary infinite-region limit, the operator determinant becomes the spatial spectral integral
\[
C
=
\frac1{2\pi}\int_{-\infty}^{\infty}
\log\Bigl(1+\frac{|G(\kappa)|^2S_J(\kappa)}{\sigma^2}\Bigr)\,d\kappa,
\]
with water-filling
\[
S_J(\kappa)
=
\Bigl(\frac1{2\pi\mu}-\frac{\sigma^2}{|G(\kappa)|^2}\Bigr)^+.
\]
This is presented as the spatial-domain analog of Shannon’s classical continuous-time AWGN result [2111.00496].

Across these formulations, the determinant plays a unifying but not uniform role. In covariance-based LD entropy it defines a mutual-information-like dependence tied to second-order structure; in regression it controls a sharp lower bound through conditional prediction error; in Gaussian-process, Gaussian-DAG, and random-field models it gives exact mutual information or conditional mutual information; and in matrix-based entropy methods it reappears as the \(\alpha\to1\) limit of spectral entropy constructions that are optimized directly from data [2205.00794, 1407.7165, 2301.08164, 2602.12346, 2606.22301, 2111.00496].

Source: https://www.emergentmind.com/topics/determinant-based-mutual-information