Papers
Topics
Authors
Recent
Search
2000 character limit reached

αβ-log-det Divergence Overview

Updated 2 March 2026
  • αβ-log-det divergence is a two-parameter family defining information divergences for SPD matrices, unifying and interpolating classical measures like Stein’s loss and AIRM.
  • It exhibits strong geometric traits such as affine invariance, nonnegativity, spectral separability, and smoothness, with extensions to infinite-dimensional positive-definite operators.
  • Its flexible parameter mapping allows precise tuning for diverse applications in machine learning, signal processing, and multiway covariance modeling.

The αβ-log-det divergence (also known as the Alpha-Beta Log-Determinant or ABLD divergence) is a two-parameter family of information divergences between symmetric positive definite (SPD) matrices and admits generalizations to positive-definite operators on infinite-dimensional Hilbert spaces. The αβ-log-det divergence parametrically unifies and interpolates between well-known matrix divergences, including Stein’s loss, Jensen–Bregman LogDet (JBLD) divergence, Bhattacharyya (log-det zero) divergence, and the affine-invariant Riemannian distance (AIRM). Its parametric flexibility, strong geometric and invariance properties, and generalizability to the infinite-dimensional setting make it a central object in information geometry, statistical manifold analysis, and applications involving SPD representations in machine learning and signal processing (Cichocki et al., 2014, Cherian et al., 2021, Quang, 2017, Quang, 2016).

1. General Definition and Domain

Let P,Q∈S+nP, Q \in S^n_+ be n×nn\times n real symmetric positive-definite matrices. For real parameters α,β\alpha, \beta satisfying α≠0\alpha \neq 0, β≠0\beta \neq 0, and α+β≠0\alpha + \beta \neq 0, the αβ-log-det divergence is defined as

DAB(α,β)(P ∥ Q)=1αβlog⁡det⁡(α(P−1Q)β+β(P−1Q)−αα+β).D^{(\alpha, \beta)}_{AB}(P \,\|\, Q) = \frac{1}{\alpha \beta} \log \det \left( \frac{\alpha (P^{-1}Q)^{\beta} + \beta (P^{-1}Q)^{-\alpha}}{\alpha + \beta} \right).

Letting λ1,…,λn\lambda_1,\ldots,\lambda_n be the eigenvalues of M=P−1QM = P^{-1}Q, an equivalent spectral form is

DAB(α,β)(P ∥ Q)=1αβ∑i=1nlog⁡(α λiβ+β λi−αα+β).D^{(\alpha, \beta)}_{AB}(P \,\|\, Q) = \frac{1}{\alpha \beta} \sum_{i=1}^n \log \left( \frac{\alpha\,\lambda_i^{\beta} + \beta\,\lambda_i^{-\alpha}}{\alpha + \beta} \right).

This function admits continuous extensions for boundary cases (n×nn\times n0, n×nn\times n1, n×nn\times n2) via L’Hôpital’s rule, yielding affine-invariant Riemannian metrics in the n×nn\times n3 limit, and other forms for degenerate parameter combinations (Cichocki et al., 2014).

The domain of validity depends on the signs of n×nn\times n4 and n×nn\times n5. For n×nn\times n6, the divergence is always finite. For n×nn\times n7, additional eigenvalue constraints are required: e.g., n×nn\times n8 for n×nn\times n9.

2. Special Cases and Recovery of Classical Divergences

By tuning α,β\alpha, \beta0, the αβ-log-det divergence specializes to many classical divergences:

α,β\alpha, \beta1 Divergence (Name) Formula / Property
α,β\alpha, \beta2 or α,β\alpha, \beta3 Stein’s loss / Burg divergence α,β\alpha, \beta4
α,β\alpha, \beta5 α,β\alpha, \beta6-log-det divergence α,β\alpha, \beta7
α,β\alpha, \beta8 Power log-det α,β\alpha, \beta9
α≠0\alpha \neq 00 S-divergence (JBLD) α≠0\alpha \neq 01
α≠0\alpha \neq 02 Affine-invariant Riemannian metric (AIRM) α≠0\alpha \neq 03
α≠0\alpha \neq 04 or α≠0\alpha \neq 05 sym. Jeffreys-KL α≠0\alpha \neq 06

Thus, the αβ-family comprises Burg (Stein’s) divergence, JBLD, Bhattacharyya distance (α≠0\alpha \neq 07), symmetrized KL, AIRM, Cauchy-Schwarz, and other divergences as instances or limits (Cichocki et al., 2014, Cherian et al., 2021, Quang, 2016).

3. Key Properties and Structure: α≠0\alpha \neq 08 Parameter Map

The repertoire of divergences covered by the αβ-log-det family is well captured by examining the α≠0\alpha \neq 09-plane:

  • β≠0\beta \neq 00 or β≠0\beta \neq 01: generalized Stein's loss forms
  • β≠0\beta \neq 02: symmetric power-log-det divergences
  • β≠0\beta \neq 03: β≠0\beta \neq 04-log-det divergences (relating to β≠0\beta \neq 05-connections in information geometry)
  • β≠0\beta \neq 06: S-divergence (JBLD)
  • β≠0\beta \neq 07: AIRM

Properties of β≠0\beta \neq 08 (Cichocki et al., 2014):

  • Nonnegativity: β≠0\beta \neq 09, with equality iff α+β≠0\alpha + \beta \neq 00
  • Definiteness: α+β≠0\alpha + \beta \neq 01
  • Smoothness: in α+β≠0\alpha + \beta \neq 02 and α+β≠0\alpha + \beta \neq 03 except for removable singularities
  • Spectral Separability: function of eigenvalues α+β≠0\alpha + \beta \neq 04, i.e., α+β≠0\alpha + \beta \neq 05
  • Affine-invariance: α+β≠0\alpha + \beta \neq 06 for invertible α+β≠0\alpha + \beta \neq 07
  • Scale-invariance: α+β≠0\alpha + \beta \neq 08, α+β≠0\alpha + \beta \neq 09
  • Inversion Duality: DAB(α,β)(P ∥ Q)=1αβlog⁡det⁡(α(P−1Q)β+β(P−1Q)−αα+β).D^{(\alpha, \beta)}_{AB}(P \,\|\, Q) = \frac{1}{\alpha \beta} \log \det \left( \frac{\alpha (P^{-1}Q)^{\beta} + \beta (P^{-1}Q)^{-\alpha}}{\alpha + \beta} \right).0
  • Dual Symmetry: DAB(α,β)(P ∥ Q)=1αβlog⁡det⁡(α(P−1Q)β+β(P−1Q)−αα+β).D^{(\alpha, \beta)}_{AB}(P \,\|\, Q) = \frac{1}{\alpha \beta} \log \det \left( \frac{\alpha (P^{-1}Q)^{\beta} + \beta (P^{-1}Q)^{-\alpha}}{\alpha + \beta} \right).1
  • On Diagonal ⟹ Metric: For DAB(α,β)(P ∥ Q)=1αβlog⁡det⁡(α(P−1Q)β+β(P−1Q)−αα+β).D^{(\alpha, \beta)}_{AB}(P \,\|\, Q) = \frac{1}{\alpha \beta} \log \det \left( \frac{\alpha (P^{-1}Q)^{\beta} + \beta (P^{-1}Q)^{-\alpha}}{\alpha + \beta} \right).2, DAB(α,β)(P ∥ Q)=1αβlog⁡det⁡(α(P−1Q)β+β(P−1Q)−αα+β).D^{(\alpha, \beta)}_{AB}(P \,\|\, Q) = \frac{1}{\alpha \beta} \log \det \left( \frac{\alpha (P^{-1}Q)^{\beta} + \beta (P^{-1}Q)^{-\alpha}}{\alpha + \beta} \right).3 satisfies the triangle inequality

This suggests that by moving along key lines or points in parameter space, practitioners can target divergences appropriate for a given problem, modulating sensitivity to spectrum or volume (Cichocki et al., 2014).

4. Infinite-Dimensional and Operator Extensions

The αβ-log-det divergence extends naturally from finite-dimensional SPD matrices to infinite-dimensional positive-definite operators, notably unitized trace-class and Hilbert-Schmidt operators on separable Hilbert spaces (Quang, 2017, Quang, 2016). In this context, extensions of the determinant—the Fredholm determinant for trace-class and the Hilbert–Carleman determinant for Hilbert–Schmidt perturbations—enable well-defined divergence formulas.

Infinite-dimensional Alpha-Beta Log-Det divergences take the form

DAB(α,β)(P ∥ Q)=1αβlog⁡det⁡(α(P−1Q)β+β(P−1Q)−αα+β).D^{(\alpha, \beta)}_{AB}(P \,\|\, Q) = \frac{1}{\alpha \beta} \log \det \left( \frac{\alpha (P^{-1}Q)^{\beta} + \beta (P^{-1}Q)^{-\alpha}}{\alpha + \beta} \right).4

where DAB(α,β)(P ∥ Q)=1αβlog⁡det⁡(α(P−1Q)β+β(P−1Q)−αα+β).D^{(\alpha, \beta)}_{AB}(P \,\|\, Q) = \frac{1}{\alpha \beta} \log \det \left( \frac{\alpha (P^{-1}Q)^{\beta} + \beta (P^{-1}Q)^{-\alpha}}{\alpha + \beta} \right).5 denotes the appropriate extended determinant, and DAB(α,β)(P ∥ Q)=1αβlog⁡det⁡(α(P−1Q)β+β(P−1Q)−αα+β).D^{(\alpha, \beta)}_{AB}(P \,\|\, Q) = \frac{1}{\alpha \beta} \log \det \left( \frac{\alpha (P^{-1}Q)^{\beta} + \beta (P^{-1}Q)^{-\alpha}}{\alpha + \beta} \right).6 are positive-definite unitized operators. Limits DAB(α,β)(P ∥ Q)=1αβlog⁡det⁡(α(P−1Q)β+β(P−1Q)−αα+β).D^{(\alpha, \beta)}_{AB}(P \,\|\, Q) = \frac{1}{\alpha \beta} \log \det \left( \frac{\alpha (P^{-1}Q)^{\beta} + \beta (P^{-1}Q)^{-\alpha}}{\alpha + \beta} \right).7, DAB(α,β)(P ∥ Q)=1αβlog⁡det⁡(α(P−1Q)β+β(P−1Q)−αα+β).D^{(\alpha, \beta)}_{AB}(P \,\|\, Q) = \frac{1}{\alpha \beta} \log \det \left( \frac{\alpha (P^{-1}Q)^{\beta} + \beta (P^{-1}Q)^{-\alpha}}{\alpha + \beta} \right).8 recover the infinite-dimensional AIRM, while DAB(α,β)(P ∥ Q)=1αβlog⁡det⁡(α(P−1Q)β+β(P−1Q)−αα+β).D^{(\alpha, \beta)}_{AB}(P \,\|\, Q) = \frac{1}{\alpha \beta} \log \det \left( \frac{\alpha (P^{-1}Q)^{\beta} + \beta (P^{-1}Q)^{-\alpha}}{\alpha + \beta} \right).9 yields the infinite-dimensional Stein divergence (Quang, 2017, Quang, 2016).

For Regularized Kernel covariance operators (λ1,…,λn\lambda_1,\ldots,\lambda_n0, λ1,…,λn\lambda_1,\ldots,\lambda_n1) in a Reproducing Kernel Hilbert Space (RKHS), the αβ-log-det divergence reduces to a Gram-matrix formula: λ1,…,λn\lambda_1,\ldots,\lambda_n2 providing a computational path for infinite-dimensional divergences in learning applications (Quang, 2016).

5. Connections to Gaussian and Information Geometric Divergences

For multivariate normal densities λ1,…,λn\lambda_1,\ldots,\lambda_n3, λ1,…,λn\lambda_1,\ldots,\lambda_n4, the continuous gamma divergence λ1,…,λn\lambda_1,\ldots,\lambda_n5 is directly expressible via the αβ-log-det divergence: λ1,…,λn\lambda_1,\ldots,\lambda_n6 with λ1,…,λn\lambda_1,\ldots,\lambda_n7 (Cichocki et al., 2014).

Special cases:

  • λ1,…,λn\lambda_1,\ldots,\lambda_n8: Kullback-Leibler divergence
  • λ1,…,λn\lambda_1,\ldots,\lambda_n9: Bhattacharyya distance
  • M=P−1QM = P^{-1}Q0: Rényi divergence of order M=P−1QM = P^{-1}Q1
  • M=P−1QM = P^{-1}Q2: Cauchy-Schwarz divergence

This reveals that αβ-log-det divergences not only cover matrix-level divergences but also bridge to statistical divergences between distributions.

6. Symmetrizations and Metric Properties

The αβ-log-det divergence is asymmetric in general. Two canonical symmetrizations are employed (Cichocki et al., 2014):

  1. Type-1 (Jeffreys-style):

M=P−1QM = P^{-1}Q3

  1. Type-2 (Jensen–Shannon style):

M=P−1QM = P^{-1}Q4

Type-1 subsumes the Jeffreys-KL divergence (when M=P−1QM = P^{-1}Q5 or vice versa), and is symmetric when M=P−1QM = P^{-1}Q6. For M=P−1QM = P^{-1}Q7, the square root of the divergence yields the affine-invariant Riemannian metric, satisfying the triangle inequality, and thus endowing the space of SPD matrices with a geodesic structure (Cichocki et al., 2014).

7. Applications, Learning, and Multiway Extensions

Recent work exploits the αβ-log-det divergence as a learnable meta-divergence for applications requiring similarity assessment between SPD matrices (Cherian et al., 2021). In supervised and unsupervised tasks (e.g., discriminative dictionary learning, clustering), parameters M=P−1QM = P^{-1}Q8—even allowed to be vector-valued—are optimized jointly with SPD dictionaries/centroids via Riemannian optimization schemes, harnessing the flexibility of the divergence family.

Empirical evaluation on multiple vision benchmarks demonstrates the advantage of automatically selecting from the αβ-family, with per-dictionary-atom vector-valued parameters yielding further performance gains (Cherian et al., 2021).

A multiway (Kronecker-separable) extension exists for block-covariance structures. For M=P−1QM = P^{-1}Q9 with Kronecker decompositions, the divergence splits into a sum of per-mode αβ-log-det divergences and a scale term, extending Hilbert, AIRM, and Stein divergences to the multi-tensor setting (Cichocki et al., 2014):

DAB(α,β)(P ∥ Q)=1αβ∑i=1nlog⁡(α λiβ+β λi−αα+β).D^{(\alpha, \beta)}_{AB}(P \,\|\, Q) = \frac{1}{\alpha \beta} \sum_{i=1}^n \log \left( \frac{\alpha\,\lambda_i^{\beta} + \beta\,\lambda_i^{-\alpha}}{\alpha + \beta} \right).0

This suggests a natural fit for multi-modal and tensor-valued covariance modeling, especially for multiway Gaussian models or tensor factor analysis.


The αβ-log-det divergence family provides a unified, parameterized, and geometrically well-motivated divergence for SPD matrices and operators, enabling fine control over spectrum sensitivity and metric properties, with theoretical guarantees and proven utility in information geometry, statistical learning, and high-dimensional covariance modeling (Cichocki et al., 2014, Cherian et al., 2021, Quang, 2017, Quang, 2016).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to αβ-log-det Divergence.