---
title: Power-Deformed Bures–Wasserstein Metric
url: https://www.emergentmind.com/topics/power-deformed-generalized-bures-wasserstein-metric
type: topic
---

# Power-Deformed Bures–Wasserstein Metric

The power-deformed generalized Bures–Wasserstein metric denotes a family of constructions that deform classical Bures–Wasserstein geometry by a power parameter while preserving a Riemannian-geometric interpretation. In the quantum-state setting, the deformation is realized by changing the base point at which the Bures–Wasserstein manifold is linearized, producing generalized fidelities and generalized Bures distances that interpolate among Uhlmann, Holevo, and Matsumoto fidelities, and that extend naturally to Rényi-type divergences [2410.04937]. In the symmetric positive definite (SPD) setting, the deformation is realized by combining a generalized Bures–Wasserstein metric parameterized by an SPD matrix with a matrix-power diffeomorphism, yielding a pullback metric used in representation learning and Riemannian batch normalization [2504.00660]. These constructions sit within a broader line of work that relates Bures–Wasserstein geometry to Gaussian optimal transport, generalized Wasserstein costs, and matrix-valued transport metrics [2110.10464][2011.05845][1808.05064].

## 1. Classical Bures–Wasserstein geometry

Let
$$
D_d := \{\rho \in H_d : \rho \succ 0,\ \operatorname{Tr}\rho = 1\}
$$
be the manifold of density operators, and let
$$
P_d := \{P \in H_d : \lambda_i(P) > 0\}
$$
be the cone of positive definite matrices. The tangent space at $\rho \in D_d$ is
$$
T_\rho D_d = \{X \in H_d : \operatorname{Tr}X = 0\},
$$
while on $P_d$ the tangent space is $H_d$ without the trace constraint. The Bures, or symmetric logarithmic derivative, Riemannian metric is defined through the Lyapunov operator $L_\rho$, where $L_\rho(Y)$ is the unique solution of $\rho X + X\rho = Y$, by
$$
\langle X, Y\rangle_\rho^{BW}
=
\operatorname{Tr}[L_\rho(X)\rho L_\rho(Y)]
=
\operatorname{Tr}[X L_\rho^{-1}(Y)].
$$
On $P_d$, the corresponding inner product is
$$
\langle U,V\rangle_P^{BW}=\operatorname{Tr}[L_P(U) P L_P(V)].
$$
The associated Bures–Wasserstein distance is
$$
d_{BW}(P,Q)
=
\sqrt{\operatorname{Tr}P+\operatorname{Tr}Q-2\operatorname{Tr}\sqrt{\sqrt{P}Q\sqrt{P}} }.
$$
It coincides with the $2$-Wasserstein distance between zero-mean Gaussian measures when $P$ and $Q$ are covariance matrices, and the Bures–Wasserstein geodesic is the displacement interpolation in that setting [2410.04937].

The Bures–Wasserstein geometry admits closed-form geodesic, exponential, and logarithm maps. If $S:=P^{-1}\#Q$, with geometric mean
$$
A\# B = A^{1/2}\sqrt{A^{-1/2}BA^{-1/2}}\,A^{1/2},
$$
then
$$
\gamma_{PQ}^{BW}(t)=[(1-t)I+tS]\,P\,[(1-t)I+tS],
$$
$$
\operatorname{Exp}_P^{BW}[V]=(I+L_P(V))P(I+L_P(V)),
$$
and
$$
\operatorname{Log}_P^{BW}[Q]=L_P^{-1}\big((P^{-1}\#Q)-I\big).
$$
In the SPD-manifold literature, the same geometry is written with a real-symmetric Lyapunov operator $\mathcal L_X$ solving
$$
X\mathcal L_X(S)+\mathcal L_X(S)X=S,
$$
and metric tensor
$$
g_X^{BW}(S_1,S_2)=\frac12 \operatorname{tr}(\mathcal L_X(S_1)S_2),
$$
highlighting the same linear dependence on the eigenvalues of the base point that later motivates robustness to ill-conditioning [2504.00660].

## 2. Generalized fidelity and generalized Bures distance

A generalized fidelity is obtained by fixing a base point $R \in P_d$ and defining
$$
F_R(P,Q)
:=
\operatorname{Tr}\!\left[
\sqrt{R^{1/2}PR^{1/2}}\,R^{-1}\sqrt{R^{1/2}QR^{1/2}}
\right].
$$
Equivalent forms include
$$
F_R(P,Q)=\operatorname{Tr}[Q^{1/2}U_Q U_P^* P^{1/2}]
=
\operatorname{Tr}[(R^{-1}\#Q)\,R\,(R^{-1}\#P)],
$$
where $U_P=\operatorname{Pol}(P^{1/2}R^{1/2})$ and $U_Q=\operatorname{Pol}(Q^{1/2}R^{1/2})$ are unitary polar factors. The generalized squared Bures–Wasserstein distance at base $R$ is then
$$
B_R(P,Q):=\operatorname{Tr}[P+Q]-2\,\operatorname{Re}F_R(P,Q),
$$
with
$$
b_R(P,Q):=\sqrt{B_R(P,Q)}.
$$
For fixed $R \succ 0$, $b_R$ is a bona fide metric on $P_d$: it is symmetric, non-negative, vanishes only when $P=Q$, and satisfies the triangle inequality [2410.04937].

The central geometric fact is that $B_R$ is the natural distance induced by linearizing the Bures–Wasserstein manifold at $R$. The tangent-space characterization states
$$
B_R(P,Q)=\|\operatorname{Log}_R^{BW}[P]-\operatorname{Log}_R^{BW}[Q]\|_R^2,
$$
where
$$
\operatorname{Log}_R^{BW}[Q]=L_R^{-1}(R^{-1}\#Q-I),
\qquad
\langle X,Y\rangle_R^{BW}=\operatorname{Tr}[L_R(X)R L_R(Y)].
$$
An equivalent Hilbert–Schmidt characterization is
$$
B_R(P,Q)=\|U_P^*P^{1/2}-U_Q^*Q^{1/2}\|_2^2.
$$
These formulas make explicit that the generalized distance is not an ad hoc modification of fidelity, but the squared norm of a difference vector in the linearized Bures–Wasserstein tangent space [2410.04937].

Several standard fidelities are recovered by special choices of the base.

| Base choice | Recovered fidelity | Condition |
|---|---|---|
| $R=P$ or $R=Q$ | Uhlmann fidelity $F_U$ | Also along the BW geodesic segment between $P$ and $Q$ |
| $R=I$ | Holevo fidelity $F_H=\operatorname{Tr}[P^{1/2}Q^{1/2}]$ | Exact reduction |
| $R=P^{-1}$ or $R=Q^{-1}$ | Matsumoto fidelity $F_M=\operatorname{Tr}[P\#Q]$ | Also along specified AI/Euc inverse paths |

This reduction pattern is one of the defining features of the generalized construction: the base point selects which classical comparator is recovered [2410.04937].

## 3. Power deformation by moving the linearization base

The power-deformed generalized Bures–Wasserstein metric arises by moving the base point along affine-invariant geodesics connecting a state to its inverse. For $x \in \mathbb R$, one sets $R=P^x$ or $R=Q^x$ and defines
$$
F_x^{(P)}(P,Q):=F_{P^x}(P,Q), \qquad
F_x^{(Q)}(P,Q):=F_{Q^x}(P,Q),
$$
together with
$$
B_x^{(P)}(P,Q):=B_{P^x}(P,Q), \qquad
B_x^{(Q)}(P,Q):=B_{Q^x}(P,Q).
$$
Explicitly,
$$
F_{P^x}(P,Q)
=
\operatorname{Tr}\!\left[P^{(1-x)/2}\sqrt{P^{x/2}QP^{x/2}}\right]
=
\operatorname{Tr}[Q^{1/2}U_xP^{1/2}],
$$
where $U_x=\operatorname{Pol}(P^{x/2}Q^{1/2})$, and similarly
$$
F_{Q^x}(P,Q)
=
\operatorname{Tr}\!\left[Q^{(1-x)/2}\sqrt{Q^{x/2}PQ^{x/2}}\right].
$$
A symmetrized polar fidelity is
$$
\bar F_x(P,Q):=\frac12\big(F_{P^x}(P,Q)+F_{Q^x}(P,Q)\big).
$$
At $x=1,0,-1$, these formulas recover Uhlmann, Holevo, and Matsumoto fidelities, respectively [2410.04937].

A commonly used parameterization fixes a reference state $\rho_0 \succ 0$ and sets
$$
R_\alpha(\rho_0):=\rho_0^{1-2\alpha}, \qquad \alpha \in [0,1].
$$
The $\alpha$-deformed generalized fidelity and distance are
$$
F_{\alpha,\rho_0}(\rho,\sigma):=F_{R_\alpha(\rho_0)}(\rho,\sigma),
$$
$$
B_{\alpha,\rho_0}(\rho,\sigma):=B_{R_\alpha(\rho_0)}(\rho,\sigma),
\qquad
b_{\alpha,\rho_0}(\rho,\sigma):=\sqrt{B_{\alpha,\rho_0}(\rho,\sigma)}.
$$
Equivalently,
$$
B_{\alpha,\rho_0}(\rho,\sigma)
=
\|\operatorname{Log}_{R_\alpha}^{BW}[\rho]-\operatorname{Log}_{R_\alpha}^{BW}[\sigma]\|_{R_\alpha}^2.
$$
The special values are
$$
\alpha=0 \Rightarrow R_0=\rho_0,\qquad
\alpha=\tfrac12 \Rightarrow R_{1/2}=I,\qquad
\alpha=1 \Rightarrow R_1=\rho_0^{-1}.
$$
Thus $\alpha=\tfrac12$ yields Holevo fidelity directly, while the endpoints reproduce Uhlmann-type and Matsumoto-type behavior under the corresponding geodesic conditions [2410.04937].

The same source relates this deformation to monotone quantum Fisher-information metrics. The Bures metric corresponds to the symmetric logarithmic derivative, the right logarithmic derivative is linked to Matsumoto’s fidelity, and Holevo-related quantities lie between them. The path
$$
R_\alpha(\rho_0)=\rho_0^{1-2\alpha}
$$
therefore provides a tractable interpolation from SLD/Bures at $\alpha=0$, through Holevo at $\alpha=\tfrac12$, to RLD/Matsumoto at $\alpha=1$. Numerically, along the AI-power paths $R=P^{1-2\alpha}$ and $R=Q^{1-2\alpha}$, generalized fidelity appears monotone in $\alpha$ and bounded between Matsumoto and Uhlmann fidelities, in line with the inequalities $F_M \le F_H \le F_U$ [2410.04937].

## 4. Path invariance, purification, and divergence extensions

The generalized fidelity is not arbitrary in its dependence on the base. When the base lies on specific geodesic-related paths, the generalized quantity collapses to familiar fidelities. If $R$ lies on the Bures–Wasserstein geodesic $\gamma_{PQ}^{BW}(t)$ between $P$ and $Q$, or on the inverse of the Bures–Wasserstein geodesic between $P^{-1}$ and $Q^{-1}$, then
$$
F_R(P,Q)=F_U(P,Q)
$$
for all $t \in [0,1]$. If $R$ lies on certain affine-invariant or Euclidean geodesic paths between inverses, or on the inverse of a Euclidean geodesic between $P$ and $Q$, then
$$
F_R(P,Q)=F_M(P,Q).
$$
There are also covariance identities for mixed Bures–Wasserstein paths, such as
$$
F_{\gamma_{P Q^{-1}}^{BW}(t)}(P,Q)=F_{\gamma_{Q P^{-1}}^{BW}(t)}(P,Q).
$$
These results formalize the statement that changing the linearization point can preserve, rather than alter, the comparator when the motion of the base is synchronized with the geometry of $P$ and $Q$ [2410.04937].

The generalized fidelity also admits an SDP-flavored block-matrix characterization. For the primal SDP with block-diagonal constraint matrix $B=\operatorname{diag}(P,Q,R)$ and optimal Gram matrix
$$
X_* = TT^*,\qquad
T=
\begin{bmatrix}
P^{1/2}U_P\\
Q^{1/2}U_Q\\
R^{1/2}
\end{bmatrix},
$$
one obtains
$$
F_R(P,Q)=\langle K,X_*\rangle,\qquad
\operatorname{Re}F_R(P,Q)=\left\langle \frac{K+K^*}{2},X_*\right\rangle,
$$
and
$$
B_R(P,Q)=\langle J,X_*\rangle.
$$
An Uhlmann-like theorem complements this representation: if $|\Omega\rangle=\sum_i |i,i\rangle$ and
$$
|P_{U_P^\top}\rangle=(P^{1/2}\otimes U_P^\top)|\Omega\rangle,\qquad
|Q_{U_Q^\top}\rangle=(Q^{1/2}\otimes U_Q^\top)|\Omega\rangle,
$$
then
$$
F_R(P,Q)=\langle P_{U_P^\top},Q_{U_Q^\top}\rangle.
$$
Generalized fidelity is therefore a carefully chosen purification overlap, extending the role played by Uhlmann’s theorem [2410.04937].

The same framework extends to multivariate fidelities and Rényi divergences. For states $\{\rho_i\}_{i=1}^n$ and base $\sigma$, the generalized multivariate fidelity is
$$
F_\sigma(\rho_1,\dots,\rho_n)
=
\frac{1}{n(n-1)}\sum_{i\ne j}F_\sigma(\rho_i,\rho_j).
$$
If $\sigma$ maximizes $\sum_i F_U(\rho_i,\sigma)$, equivalently is the Bures–Wasserstein barycenter up to normalization, then
$$
F_\sigma(\rho_1,\dots,\rho_n)=\frac{f(\sigma)^2-n}{n^2-n}.
$$
For Rényi-type quantities, the base-dependent trace functional
$$
F_R^\alpha(\rho,\sigma)
:=
\operatorname{Tr}\!\left[
(R^{1/2}\rho R^{1/2})^\alpha
R^{-1}
(R^{1/2}\sigma R^{1/2})^{1-\alpha}
\right]
$$
induces
$$
\hat D_{\alpha,R}(\rho\|\sigma)
=
\frac{1}{\alpha-1}\log \operatorname{Re}F_R^\alpha(\rho,\sigma).
$$
Specific bases recover Petz, sandwiched, reverse-sandwiched, and geometric Rényi divergences:
$$
R=I,\quad
R=\sigma^{(1-\alpha)/\alpha},\quad
R=\rho^{\alpha/(1-\alpha)},\quad
R=\sigma^{-1},
$$
respectively. In this sense, the base acts as the deformation variable that unifies several non-commutative Rényi quantizations [2410.04937].

## 5. Generalized Bures–Wasserstein metrics on the SPD manifold

A distinct but closely related formulation was introduced for SPD matrices. For $\mathcal S_{++}^d$, the generalized Bures–Wasserstein metric is parameterized by an SPD matrix $M$ and defined by
$$
g_X^{GBW}(S_1,S_2)
=
\frac12 \operatorname{tr}\!\left(\mathcal L_{X,M}(S_1)S_2\right)
=
\frac12 \operatorname{vec}(S_1)^T (X\otimes M + M\otimes X)^{-1}\operatorname{vec}(S_2),
$$
where the generalized Lyapunov operator solves
$$
X\mathcal L_{X,M}(S)M + M\mathcal L_{X,M}(S)X = S.
$$
When $M=I$, this reduces to the classical Bures–Wasserstein metric. When $M=X$, it coincides locally with the affine-invariant metric. The associated distance can be seen as the Bures–Wasserstein distance between congruence-transformed matrices, and the geometry is Riemannian-isometric to the Bures–Wasserstein geometry under the congruence map [2110.10464][2504.00660].

The power-deformed GBWM introduced for learning on SPD manifolds combines this metric with the matrix-power diffeomorphism
$$
\phi_\theta(X)=X^\theta,\qquad \theta\in\mathbb R\setminus\{0\},
$$
and defines
$$
g_X^{(\theta)-GBW}(S_1,S_2)
=
\frac{1}{\theta^2}
g_{X^\theta}^{GBW}\!\big((\phi_\theta)_{*,X}(S_1),(\phi_\theta)_{*,X}(S_2)\big).
$$
The deformation therefore acts by transporting tangent vectors through the matrix-power map and rescaling the metric by $1/\theta^2$. The limit $\theta\to 0$ is a Log-Euclidean-type, or LEM-like, metric:
$$
g_X^{(\theta)-GBW}(S,S)
\overset{\theta\to 0}{\longrightarrow}
\frac12
\langle \log_{*,X}(\mathcal L_M(S)),\log_{*,X}(S)\rangle,
$$
and locally
$$
g_X^{(\theta)-GBW}(S_1,S_2)=\frac14 g_X^{(\theta)-AI}(S_1,S_2).
$$
The deformation interpolates between GBWM at $\theta=1$ and an LEM-like regime as $\theta\to 0$ [2504.00660].

The isometric structure is central. If
$$
\varphi(X)=M^{1/2}XM^{1/2},
\qquad
f=\phi_\theta\circ\varphi,
$$
then $f$ is a Riemannian isometry from $(\mathcal S_{++}^d,g^{BW})$ to $(\mathcal S_{++}^d,g^{(\theta)-GBW})$. Consequently, exponential maps, logarithms, parallel transport, distances, and weighted Fréchet means in the deformed geometry are computed by mapping to the Bures–Wasserstein manifold, performing the operation there, and mapping back. This construction supplies the geometric basis for the deformed Riemannian batch-normalization layer used in SPD deep networks [2504.00660].

## 6. Algorithms and learning applications

For the generalized fidelity $F_R(P,Q)$ in the quantum-state formulation, a direct evaluation uses polar factors. One computes
$$
U_P=\operatorname{Pol}(P^{1/2}R^{1/2}),\qquad
U_Q=\operatorname{Pol}(Q^{1/2}R^{1/2}),
$$
for example through
$$
\operatorname{Pol}(M)=M(\sqrt{M^*M})^{-1},
$$
and then evaluates
$$
F_R(P,Q)=\operatorname{Tr}[Q^{1/2}U_Q U_P^* P^{1/2}],
$$
followed by
$$
b_R(P,Q)=\sqrt{\operatorname{Tr}[P+Q]-2\operatorname{Re}F_R(P,Q)}.
$$
Computing $\sqrt{M^*M}$ uses spectral decomposition or Schur methods, with complexity $O(d^3)$. An alternative uses Lyapunov solves to evaluate logarithm maps and tangent norms, with Bartels–Stewart or sign-function methods noted for numerical stability. In the power-deformed case, the base $R=\rho_0^{1-2\alpha}$ or $R=P^x,Q^x$ is computed by spectral decomposition [2410.04937].

In the SPD-learning formulation, the deformed geometry is implemented isometrically. For deformed GBWM batch normalization, one maps the input batch to Bures–Wasserstein coordinates via
$$
\hat X_i = M^{-1/2}X_i^\theta M^{-1/2},
\qquad
\hat{\mathcal G}=M^{-1/2}\mathcal G^\theta M^{-1/2},
$$
computes the Bures–Wasserstein batch mean and variance, performs centering, scaling, and biasing in the Bures–Wasserstein manifold, and maps back using
$$
\acute X_i = (M^{1/2}\tilde X_i M^{1/2})^{1/\theta}.
$$
The batch variance is normalized by $1/\theta^2$ because of the metric scaling. Differentiation relies on the Daleckii–Kreĭn formula for matrix powers and square roots and on explicit gradients for Lyapunov operators [2504.00660].

Empirical validation was reported on HDM05 action recognition, NTU RGB+D action recognition, and MAMEM-SSVEP-II EEG classification, with SPDNet and RResNet backbones. The reported results include $65.47\% \pm 3.33$ for SPDNet-GBWBN with $\theta=1$ on MAMEM-SSVEP-II, $69.10\% \pm 0.83$ for RResNet-GBWBN with $\theta=0.5$ on HDM05, and $59.72\% \pm 0.31$ on NTU RGB+D. The ablation on $\theta$ states that deformed metrics with $\theta\neq 1$ generally improve over $\theta=1$, and the experiments are presented as evidence that the learnable metric parameter $M$ and the power deformation improve robustness on ill-conditioned SPD matrices [2504.00660].

## 7. Related transport frameworks and open problems

The power-deformed generalized Bures–Wasserstein metric should be distinguished from broader matrix-valued transport generalizations. The generalized BW geometry of SPD matrices parameterized by $M$ was formulated earlier as a Mahalanobis-cost extension of Gaussian Bures–Wasserstein geometry, with explicit geodesics, exponential and logarithm maps, Levi-Civita connection, curvature bounds, and barycenter equations [2110.10464]. At a still broader level, weighted Wasserstein–Bures metrics on positive semidefinite matrix-valued Radon measures define complete geodesic spaces with conic structure through a convex Benamou–Brenier formulation. In that framework, a power deformation can be introduced by replacing the weight pair $\Lambda=(A_1,A_2)$ with $\Lambda_p=(A_1^p,A_2^p)$, preserving convexity, duality, geodesicity, and cone structure for each fixed $p$ [2011.05845]. The Kantorovich–Bures metric on matrix-valued measures provides another transport–reaction geometry whose constant-in-space reduction recovers the classical Bures metric, and whose structure indicates where an $\alpha$-power deformation could be inserted, although the paper itself does not define such a deformation [1808.05064].

Several open questions remain explicit in the generalized-fidelity literature. Data processing inequality and joint concavity for generalized fidelities and distances at arbitrary bases remain open; initial numerics are reported as showing no data-processing-inequality violations, but a proof is outstanding. A true SDP formulation of the generalized fidelity, more precisely of its real part, is not yet known: the current characterization is SDP-flavored rather than an SDP feasibility or optimization problem. Also open is the characterization of the image of the unitary factors $\{U_Q U_P^*\}$ over all bases $R$ as $SU(d)$ [2410.04937].

Taken together, these developments define a coherent geometric program. In the quantum-state setting, power deformation means changing the point of Bures–Wasserstein linearization through $R_\alpha=\rho_0^{1-2\alpha}$ or $R=P^x,Q^x$, thereby interpolating among named fidelities, monotone metrics, and Rényi divergences. In the SPD-learning setting, power deformation means pulling back the generalized Bures–Wasserstein metric through $X\mapsto X^\theta$, thereby producing a learnable and numerically robust geometry for normalization and optimization. A plausible implication is that the phrase “power-deformed generalized Bures–Wasserstein metric” now refers less to a single formula than to a common design principle: introduce a power parameter at the level of the geometric base or of the manifold coordinates, and use the induced pullback or linearized metric to tune the non-commutative geometry without abandoning the Bures–Wasserstein framework [2410.04937][2504.00660].

Source: https://www.emergentmind.com/topics/power-deformed-generalized-bures-wasserstein-metric