Alpha-Z Bures-Wasserstein Divergence
- Alpha-Z Bures-Wasserstein divergence is a Rényi-induced matrix divergence defined on positive definite Hermitian and symmetric PSD matrices, replacing the classical geometric term with a trace-based operator functional.
- Its formulation uses parameters α and z to control interpolation between matrices, leading to properties like nonnegativity, nonsymmetry, and a uniquely defined weighted right mean.
- In neuroimaging, the divergence is applied as a robust, geometry-aware comparison score for functional connectomes, outperforming traditional metrics in high-dimensional, rank-deficient settings.
The Alpha-Z Bures-Wasserstein divergence is a matrix divergence obtained by replacing the geometric term in the classical Bures-Wasserstein expression with the trace of an - Rényi-type operator functional. In finite-dimensional quantum matrix analysis it is defined on positive definite Hermitian matrices and treated as a nonsymmetric quantum divergence; in recent applied work it is also used on symmetric positive semidefinite matrices, especially functional connectomes, where it serves as a geometry-aware comparison score rather than as a proven geodesic metric (Jeong et al., 2022, Uddin et al., 30 Jul 2025).
1. Terminology, scope, and naming
The name refers to a specific divergence introduced from the - Rényi relative entropy, not to every parameterized generalization of Bures-Wasserstein geometry. In the finite-dimensional matrix setting, the underlying spaces are , , and , the complex matrices, Hermitian matrices, and positive definite Hermitian matrices. In applied neuroimaging, the same formula is used for real symmetric positive semidefinite matrices , because functional connectomes can be rank-deficient when the number of parcels exceeds the number of time points (Jeong et al., 2022, Uddin et al., 30 Jul 2025).
The terminology is not universal across the Bures-Wasserstein literature. Several papers on Bures-Wasserstein geometry do not define an Alpha-Z object at all. The paper on generalized fidelities and generalized Bures-Wasserstein distances introduces a base-dependent generalized fidelity, a generalized Bures distance, and a base-dependent Rényi family , but explicitly states that it does not define an object named “Alpha-0 Bures-Wasserstein Divergence” with independent 1 and 2 parameters (Afham et al., 2024). The paper on adapted optimal transport between Gaussian processes similarly states that it does not introduce an 3-divergence, 4-divergence, or any “Alpha-Z” variant of Bures-Wasserstein, but instead derives an adapted Bures-Wasserstein distance built from Cholesky factors and the diagonal of 5 (Gunasingam et al., 2024). Likewise, the Alpha Procrustes family is an 6-parameterized metric family with no 7-parameter, and the geodesic theory for Bures-Wasserstein covariance matrices of different ranks contains no 8-parametrized divergence (Quang, 2019, Thanwerdas et al., 2022).
A basic encyclopedic caution is therefore required: “Alpha-Z Bures-Wasserstein divergence” designates a particular Rényi-induced construction, whereas “generalized Bures-Wasserstein” may refer to several inequivalent programs.
2. Mathematical definition
In the quantum-divergence formulation, the central operator is
9
defined for 0 and 1. The paper describes 2 as the matrix version of the 3-4 Rényi relative entropy, and notes in particular that 5 is the sandwiched quasi-relative entropy. From this quantity, the 6-7 Bures-Wasserstein quantum divergence is defined, for 8, by
9
Equivalently,
0
This is the fundamental formal definition in the matrix-analysis literature (Jeong et al., 2022).
The neuroimaging paper uses the same construction for 1, writing
2
with
3
That paper explicitly states that the divergence is defined for positive semidefinite matrices, not only strictly positive definite ones. It also notes that the displayed formula for 4 is typographically corrupted in the PDF text, while identifying the intended expression as the one above (Uddin et al., 30 Jul 2025).
The parameters play different roles. The parameter 5 appears both in the linear trace term 6 and in the exponents inside 7, so it controls the weighting between the two arguments as well as the nonlinear operator interpolation. The parameter 8 appears in the inner exponents and the outer power, so it controls the operator-power form of the nonlinear comparison term. In the neuroimaging experiments, the reported choice is 9, for which
0
The paper calls 1 a divergence, and although it also uses phrases such as “distance measure” and “divergence-based distance metric,” it does not define a separate closed-form symmetric distance derived from 2 (Uddin et al., 30 Jul 2025).
3. Relation to classical Bures-Wasserstein geometry
The classical Bures-Wasserstein distance between positive definite matrices is commonly written as
3
Its fidelity term is
4
and the distance admits equivalent formulations such as
5
It is also the covariance-level expression of the 6-Wasserstein distance between centered Gaussian laws, and it underlies a Riemannian quotient geometry on the positive definite cone (Bhatia et al., 2017).
The Alpha-Z divergence deforms the Bures-Wasserstein formula by replacing the square-root fidelity term with the trace of 7. In the quantum-divergence paper, the special case
8
gives
9
and the paper writes
0
This expresses the divergence as a parameterized deformation of the Bures-Wasserstein squared-distance expression (Jeong et al., 2022).
Because different papers normalize the Bures-Wasserstein quantity differently, direct comparison requires attention to constants. A plausible implication is that the Alpha-Z literature is best read as preserving the arithmetic-minus-geometric structure of the Bures formula while deforming the geometric term through the 1-2 Rényi operator expression.
The classical Bures-Wasserstein framework also supplies the barycentric and geodesic background against which the Alpha-Z divergence is interpreted. For the ordinary metric, the two-point geodesic mean is
3
the midpoint is the Wasserstein mean, and the barycenter of several matrices is characterized by the fixed-point equation
4
These formulas are classical BW objects rather than Alpha-Z ones, but they are the reference point for the later right-mean theory associated with 5 (Bhatia et al., 2017).
4. Divergence properties and the associated right mean
In the quantum-divergence treatment, 6 is a quantum divergence, meaning a smooth map with nonnegativity, equality only on the diagonal, vanishing first derivative with respect to the second variable on the diagonal, and positive-semidefinite second derivative there. The paper also states that 7 in general. This nonsymmetry is structurally important because it leads to distinct right and left barycentric constructions (Jeong et al., 2022).
For a weighted tuple 8 with 9, the 0-1 weighted right mean is defined as the unique minimizer
2
Existence and uniqueness follow from strict convexity of
3
using strict concavity of 4. The minimizer is characterized as the unique positive definite solution of
5
An equivalent form is
6
where 7 denotes the weighted geometric mean (Jeong et al., 2022).
Several structural properties are known. In the commuting case,
8
so the parameter 9 disappears. The right mean is homogeneous, permutation invariant, repetition invariant, and covariant under unitary congruence. It also satisfies
0
with equality if and only if 1. If 2, then
3
The paper further proves comparisons with arithmetic means, matrix power means, and the Cartan mean, including
4
and weak log-majorization consequences involving 5 and 6 (Jeong et al., 2022).
The Wasserstein mean appears as a special case: 7 The same paper proves the trace inequality
8
This locates the Alpha-Z right mean within the classical matrix-mean hierarchy rather than outside it (Jeong et al., 2022).
In the applied paper, the divergence is described more cautiously. It states nonnegativity,
9
and reports that the divergence is invariant under completely positive trace-preserving maps, which implies the data processing inequality. It also states an in-betweenness property: 0 for any matrix power mean 1 with 2. At the same time, the paper does not establish symmetry, triangle inequality, affine invariance, or strict metric structure, and in its metric summary table it explicitly marks Alpha-Z as not geodesic (Uddin et al., 30 Jul 2025).
5. Relation to neighboring Bures-Wasserstein generalizations
The Alpha-Z divergence sits among several distinct attempts to generalize Bures-Wasserstein geometry, and these constructions should not be conflated. One line of work introduces a base-dependent generalized fidelity
3
and the associated generalized Bures distance
4
interpreted as the tangent-space distance obtained by linearizing the Bures-Wasserstein manifold at a reference point 5. That same paper introduces
6
with
7
By special choices of 8, this family recovers Petz, sandwiched, reverse-sandwiched, and geometric Rényi divergences. The paper is explicit, however, that this is not an independent two-parameter 9 Bures-Wasserstein divergence (Afham et al., 2024).
Another nearby family is the Alpha Procrustes distance
0
which yields, for 1,
2
and in the limit 3,
4
This family is a Riemannian geodesic metric family on the SPD cone, but it contains no 5-parameter and is therefore not the Alpha-Z divergence (Quang, 2019).
A different direction is the adapted Bures-Wasserstein distance arising from bicausal optimal transport for discrete-time Gaussian processes. There the covariance term becomes
6
with 7 and 8 Cholesky factors. This is filtration-sensitive and triangular rather than spectral, and the source paper explicitly states that it does not define any 9- or 00-variant (Gunasingam et al., 2024).
Finally, the ordinary Bures-Wasserstein geometry on PSD covariance matrices of arbitrary rank has its own complete geodesic theory. Minimizing geodesics between 01 and 02 are described by
03
with nonuniqueness determined by the overlap rank 04. This theory is foundational for low-rank BW geometry but does not introduce 05- or 06-deformations (Thanwerdas et al., 2022).
Taken together, these papers suggest that “generalized Bures-Wasserstein” is a broad umbrella. The Alpha-Z divergence is one specific Rényi-induced member of that broader landscape.
6. Empirical use in functional-connectome analysis
The most extensive application in the supplied literature uses the Alpha-Z Bures-Wasserstein divergence to compare functional connectomes modeled as symmetric PSD matrices. The motivation is that Pearson correlation, Euclidean distance, and common SPD-manifold geodesic distances either ignore the non-Euclidean geometry of FC data or become unstable and tuning-sensitive in high-dimensional, rank-deficient regimes. The paper therefore presents Alpha-Z as a flexible extension of Bures-Wasserstein intended to be more robust across tasks and parcellation resolutions (Uddin et al., 30 Jul 2025).
The computational use is direct. For FC matrices 07, one evaluates
08
typically by eigendecomposition-based matrix fractional powers. For the experimentally used setting 09,
10
which removes the outer matrix power. Pairwise divergences are then used in a 11-nearest-neighbor identification procedure. The reported parameter choice is fixed across experiments at
12
The paper contrasts this with affine-invariant and log-Euclidean distances, whose performance depends strongly on regularization 13 (Uddin et al., 30 Jul 2025).
The empirical evaluation covers the Human Connectome Project dataset across eight fMRI conditions—Rest, Emotion, Gambling, Language, Motor, Relational, Social, and Working Memory—and parcellation resolutions from 14 to 15 parcels. The paper states that Alpha-Z “consistently provides superior performance across all tasks and parcellation resolutions,” that it remains strong as resolution increases, and that median identification rises from about 16 at 17 parcels to 18 by 19 parcels and plateaus near 20 at 21 parcels. It further reports that AI and log-Euclidean degrade markedly in the rank-deficient regime, BW remains weaker, and Alpha-Z and Alpha-Procrustes remain robust (Uddin et al., 30 Jul 2025).
The same study also interprets the divergence in terms of subject-specific and network-specific brain fingerprints. It reports that default mode and frontoparietal control networks are especially informative in resting state, that Alpha-Z produces tighter same-subject clusters than AI in a 22D visualization at 23 parcels, and that a null model with permuted subject labels yields chance-level accuracy while Alpha-Z remains much higher. The summary states that the method offers enhanced sensitivity in linking FC patterns to cognitive and behavioral outcomes, but the paper does not present a direct predictive model of phenotype or behavior (Uddin et al., 30 Jul 2025).
In this applied setting, the Alpha-Z Bures-Wasserstein divergence functions as a geometry-aware PSD-matrix comparison score rather than as a geodesic metric. This suggests a division of labor across the literature: matrix-analysis papers emphasize variational structure, nonsymmetry, and induced means, whereas application papers emphasize robustness on semidefinite data and performance in high-dimensional comparison tasks.