---
title: 'Relative Divergence: A Cross-Disciplinary Overview'
url: https://www.emergentmind.com/topics/relative-divergence-rd
type: topic
---

# Relative Divergence: A Cross-Disciplinary Overview

Searching arXiv for recent and foundational papers on “Relative Divergence” across the senses represented in the source material.
I’ll synthesize the term as a cross-disciplinary concept, since the source material shows that “Relative Divergence (RD)” is used in multiple non-equivalent ways rather than as a single standardized definition.
“Relative Divergence” (RD) is not a single universally standardized object across the arXiv literature. The term is used for several non-equivalent constructions that compare one object to another relative to a reference structure: probability distributions relative to a baseline distribution, quantum states relative to a second state, grading functions relative to another grading on an ordered set, finitely generated groups relative to a subgroup, and model-dependent discrepancies defined through a learning procedure. This multiplicity of usage suggests that RD is best understood as a family resemblance term: a divergence-like quantity whose precise meaning is fixed by the ambient category and by the operational question under study [1206.2459] [2510.04314] [1406.4232].

## 1. Terminological scope and recurrent structural pattern

Across the literature, RD typically compares two objects by measuring how one departs from another after fixing a reference geometry, order, support condition, or hypothesis class. In some papers the object is a standard divergence between distributions; in others it is a quadratic discrepancy, a support-sensitive Rényi-type quantity, or a large-scale geometric invariant. The same abbreviation also collides with unrelated usages such as “property RD” for rapid decay, which is a representation-theoretic notion rather than a divergence [1305.0480].

| Domain | Object compared | Representative construction |
|---|---|---|
| Information theory and statistics | \(P\) relative to \(Q\) | Rényi divergence, KL divergence, relative Pearson divergence, relative extropy |
| Quantum information | \(\rho\) relative to \(\sigma\) | Petz, sandwiched, maximal, minimal Rényi-type divergences |
| Ordered structures | grading function \(F\) relative to \(G\) | increment-based RD on chains and posets |
| Geometric group theory | group \(G\) relative to subgroup \(H\) | upper and lower relative divergence |
| Model-oriented ML | dataset/distribution \(p\) relative to \(q\) through a model | R-divergence, RDR, RADAR components |

A recurring pattern is that the divergence is not merely a pointwise comparison. It is often induced by an auxiliary structure: a likelihood ratio, a support projector, a partial order, a subgroup neighborhood, a target energy, or a hypothesis minimizing empirical risk. This structural dependence is explicit in the grading-function, group-theoretic, and model-oriented formulations [2510.04314] [1406.4232] [2310.01109].

## 2. Relative divergence in probability, information theory, and density-ratio methods

In the classical information-theoretic sense, RD is frequently identified either with Kullback–Leibler divergence, also called relative entropy, or with the broader Rényi divergence family. Rényi divergence of order \(\alpha\) is defined by
\[
D_{\alpha}(P\Vert Q)=\frac{1}{\alpha-1}\ln\int p^{\alpha}q^{1-\alpha}\,d\mu,
\]
with continuous extensions at \(\alpha=0,1,\infty\). The order-\(1\) case is the KL divergence
\[
D(P\Vert Q)=\int p\log\frac{p}{q}\,d\mu,
\]
and the family is nondecreasing in \(\alpha\), satisfies data processing, and interpolates between support-sensitive and worst-case notions of distinguishability [1206.2459]. The same family is presented as a broad notion of “relative divergence” in another review, where \(D_1(P\|Q)=D(P\|Q)\), \(D_0(P\|Q)=-\ln Q(p>0)\), and \(D_\infty(P\|Q)\) is the logarithm of the essential supremum of \(dP/dQ\) [1001.4448].

A separate statistical line develops relative divergences through density-ratio smoothing. Relative density-ratio estimation replaces the ordinary ratio
\[
r(x)=\frac{p(x)}{p'(x)}
\]
by the \(\alpha\)-relative ratio
\[
r_\alpha(x)=\frac{p(x)}{\alpha p(x)+(1-\alpha)p'(x)},
\]
where the denominator is the \(\alpha\)-mixture density \(q_\alpha(x)=\alpha p(x)+(1-\alpha)p'(x)\). The associated relative Pearson divergence is
\[
\mathrm{PE}_\alpha[p,p'] =\frac{1}{2}\int \left(r_\alpha(x)-1\right)^2 q_\alpha(x)\,dx.
\]
For \(\alpha>0\), \(r_\alpha(x)\le 1/\alpha\), so the target ratio is uniformly bounded; this is the paper’s main robustness argument. The resulting RuLSIF estimator uses a least-squares objective and yields a closed-form solution \(\widehat\theta=(\widehat H+\lambda I_n)^{-1}\widehat h\) for kernel models [1106.4729].

Related relative constructions appear in more recent generative-model evaluation. There the relative density ratio is
\[
r(x)=\frac{p(x)}{\frac{1}{2}\{p(x)+q(x)\}}=\frac{2p(x)}{p(x)+q(x)},
\]
which has bounded image \([0,2]\). The paper couples this functional object to the scalar score
\[
H^2\!\left(P,\frac{P+Q}{2}\right),
\]
and proves that one-to-one transforms of the ordinary density ratio preserve \(\phi\)-divergence under pushforward, so the RDR acts as a divergence-preserving one-dimensional summary of distributional discrepancy [2510.25507].

Another extropy-based branch defines relative extropy by the quadratic expression
\[
d(f,g)=\frac{1}{2}\int_0^\infty \bigl(f(x)-g(x)\bigr)^2\,dx,
\]
together with directed extropy divergences \(J(f\mid g)\) and \(J(g\mid f)\) satisfying
\[
d(f,g)=J(f\mid g)+J(g\mid f).
\]
This is positioned as an extropy-side analogue of the relation between inaccuracy and KL divergence, and is extended to residual and past lifetime settings through conditional densities [2503.07123].

These formulations share a common design choice: direct comparison to \(Q\) is replaced or supplemented by comparison to a smoothed, transformed, or structurally induced reference. This suggests that “relative” in RD often indicates regularization of the denominator or embedding of the comparison into a more stable functional setting.

## 3. Quantum relative divergences and Rényi-type endpoint structure

In quantum information, RD usually refers to a Rényi-type divergence between positive semidefinite operators or density matrices. A central distinction is between the traditional or Petz-type quantity
\[
\widetilde D_\alpha(\rho\|\sigma)=\frac{1}{\alpha-1}\log \operatorname{Tr}(\rho^\alpha \sigma^{1-\alpha}),
\]
and the sandwiched quantum Rényi divergence
\[
D_\alpha(\rho\|\sigma):=\frac{1}{\alpha-1}\log \operatorname{Tr}\!\left( \left( \sigma^{\frac{1-\alpha}{2\alpha}} \rho \sigma^{\frac{1-\alpha}{2\alpha}} \right)^\alpha \right).
\]
For \(\alpha\ge \tfrac12\), the sandwiched divergence satisfies the data-processing inequality and recovers several operationally important quantities: \(D(\rho\|\sigma)\) as \(\alpha\to1\), \(D_{\min}\) at \(\alpha=\tfrac12\), and \(D_{\max}\) as \(\alpha\to\infty\) [1308.5961].

A key endpoint result concerns the \(0\)-relative Rényi entropy
\[
\widetilde D_0(\rho\|\sigma)=-\log\!\bigl(\operatorname{Tr}(\Pi_\rho \sigma)\bigr).
\]
Datta and Leditzky proved that
\[
\lim_{\alpha\to 0} D_\alpha(\rho\|\sigma)=\widetilde D_0(\rho\|\sigma)
\]
holds when \(\operatorname{supp}\rho=\operatorname{supp}\sigma\), but can fail if \(\operatorname{supp}\rho\subset \operatorname{supp}\sigma\) is strict. Their conclusion is that the sandwiched divergence is not by itself a universal parent quantity for all operationally relevant relative entropies: one needs the sandwiched family for \(\alpha\ge \tfrac12\) and the traditional Rényi relative entropy for \(\alpha\in[0,1)\) [1308.5961].

Another quantum line studies sufficiency of Rényi divergences for state-pair equivalence. Classically, equality of \(D_\alpha\) on an open interval determines interconvertibility of dichotomies. Quantum mechanically, known Rényi families are invariant under anti-unitary transformations, so no such family can be sufficient for CPTP-interconvertibility. The paper shows that the Petz and maximal quantum Rényi divergences remain insufficient even if one enlarges the notion of convertibility to positive trace-preserving maps, while giving evidence that the minimal or sandwiched family may be sufficient in that enlarged sense [2304.12989].

There is also a geometric equivalence result between relative \(\alpha\)-entropy and Rényi divergence. Under the escort transformation
\[
p^{(\alpha)}=\frac{p^\alpha}{\int p^\alpha\,d\mu},
\]
the paper proves
\[
\mathscr I_\alpha(P,Q)=D_{1/\alpha}\!\left(P^{(\alpha)}\|Q^{(\alpha)}\right),
\]
and shows that projection theorems, Pythagorean inequalities, associated statistical families, and divergence-induced Riemannian metrics become equivalent under this correspondence [1701.06347].

## 4. Relative divergence on ordered sets and grading functions

A distinct meaning of RD appears in the theory of grading functions on chains and posets. If \(W\) is a totally ordered chain and \(F,G\) are grading functions with increments
\[
f_k=\Delta_k F,\qquad g_k=\Delta_k G,
\]
then the relative divergence of \(F\) from \(G\) on \(W\) is
\[
\mathcal{D}(F \Vert G)\vert_W = -\sum_{k\in Z} \ln\!\left(\frac{f_k}{g_k}\right) f_k.
\]
When \(G\) is the index function with unit increments, this reduces to
\[
\mathcal{D}(F\Vert I)\vert_W = -\sum_{k\in Z} f_k \ln f_k,
\]
so Shannon entropy appears as a special case of RD [2510.04314].

The same literature treats RD as inherently dependent on order structure rather than as a universal distributional divergence. On posets, the definition is assembled from chainwise divergences using structural rules such as block-additivity on block-chains,
\[
\mathcal{D}(F\Vert G)\vert_W = \sum_{k=1}^K \mathcal{D}(F\Vert G)\vert_{W_k},
\]
and chain-infinum on even-sided split-chains,
\[
\mathcal{D}(F\Vert G)\vert_C = \inf_{\forall l}\mathcal{D}(F\Vert G)\vert_{C_l}.
\]
The paper explicitly states that there may be no universal RD definition for arbitrary posets; the definition must reflect the poset structure and the interdependence of maximal chains [2510.04314].

This framework supports the Maximum Relative Divergence Principle (MRDP): among admissible grading functions with a given null grading \(G\), choose the one maximizing RD from \(G\). On finite chains with interpolation constraints, the maximizing grading is piecewise linear:
\[
F(i)=a_k+b_k i,\qquad i\in I_k,
\]
with
\[
b_k=\frac{\Delta_k m}{\Delta_k n},\qquad a_k=m_k-b_k n_{k-1}.
\]
MRDP is then used to recover classical formulas such as
\[
P(B\mid A)=\frac{P(A\cap B)}{P(A)}
\]
and
\[
P(A\cap B)=P(A)P(B)
\]
from optimization on conjoined event posets [2510.04314].

An earlier paper develops the same program specifically for direct products of chains. There RD on a chain bundle is defined by taking the minimum over maximal chains,
\[
D(F\|G)\big|_W := \min_{MC} D(F\|G)\big|_{MC},
\]
with the natural grading \(N(\mathbf i)=i_1+\cdots+i_R\). For additively separable grading functions, RD decomposes as a sum of componentwise terms:
\[
D(F\|G)\big|_W = \sum_{r=1}^R D(F_r\|G_r)\big|_{W_r}.
\]
This makes MRDP a direct generalization of the maximum-entropy principle to ordered multidimensional state spaces [2303.14261].

## 5. Relative divergence in geometric group theory

In geometric group theory, “relative divergence” has a different meaning again. For a geodesic metric space \(X\) and subspace \(A\subset X\), one defines upper and lower relative divergence by measuring the difficulty of connecting points while avoiding neighborhoods of \(A\). With
\[
N_r(A)= \{x \in X\mid d_X(x, A)<r\},\qquad
\partial N_r(A)= \{x \in X\mid d_X(x, A)=r\},
\]
and \(d_{r,A}\) the induced length metric on \(X-N_r(A)\), the upper relative divergence is the family
\[
\delta^n_{\rho}(r)=\sup d_{\rho r}(x_1,x_2),
\]
taken over \(x_1,x_2\in \partial N_r(A)\) satisfying \(d_r(x_1,x_2)<\infty\) and \(d(x_1,x_2)\le nr\). The lower relative divergence is
\[
\sigma^n_{\rho}(r)=\inf d_{\rho r}(x_1,x_2),
\]
over points on \(\partial N_r(A)\) that are far apart in ambient distance but connected outside \(N_r(A)\) [1406.4232].

For a finitely generated group \(G\) and subgroup \(H\), these become \(Div(G,H)\) and \(div(G,H)\), quasi-isometry invariants of the pair \((G,H)\). Upper RD generalizes Gersten’s divergence, and lower RD generalizes the lower divergence of Cooper–Mihalik. A major contrast emphasized in the paper is that classical lower divergence of a one-ended finitely generated group is only linear or exponential, whereas relative lower divergence can be any polynomial degree or exponential [1406.4232].

The paper derives comparison theorems. For a finitely generated normal subgroup \(H\triangleleft G\) with one-ended quotient,
\[
Div(G/H,e)\preceq Div(G,H)\preceq Dist^H_G\circ Div(G/H,e),
\]
where \(Dist^H_G\) is upper distortion. For cyclic subgroups \(H=\langle h\rangle\), lower RD is controlled by the divergence of the subgroup axis and the subgroup distortion:
\[
div_\alpha\preceq div(G,H)\preceq div_\alpha\circ (3Dist^H_G).
\]
The theory is then applied to CAT(0) groups and relatively hyperbolic groups, showing, for example, that for any polynomial or exponential function \(f\), there exists a CAT(0) pair \((G,H)\) with \(H\cong\mathbb Z\) such that \(div(G,H)\sim f\) [1406.4232].

This version of RD is not a divergence between measures at all. It is a large-scale geometric invariant measuring how the complement of a subgroup neighborhood stretches. Its inclusion under the same label illustrates how broad the term “relative divergence” has become across fields.

## 6. Model-oriented and recent machine-learning variants

Recent machine-learning work uses “divergence” language in explicitly model-dependent ways. One example is **R-divergence**, a model-oriented discrepancy between distributions \(p\) and \(q\), defined using the hypothesis \(h_u^*\) that minimizes risk on the mixture \(u=(p+q)/2\):
\[
\mathcal D_{\mathrm R}(p\|q)=\left|\epsilon_p(h_u^*)-\epsilon_q(h_u^*)\right|.
\]
Its empirical estimator trains on pooled data \(\widehat u=\widehat p\cup\widehat q\) and compares empirical risks on the two parts. The paper emphasizes that this is not a universal \(f\)-divergence; it depends on the hypothesis class, loss, and target function, so it measures whether two datasets are effectively the same for a specific learning model [2310.01109].

Another use appears in reinforcement learning, where **relative Pearson divergence** regularizes policy updates. With importance ratio
\[
\rho=\frac{\pi}{\pi_b},
\]
mixture policy
\[
\pi_\beta=\beta\pi+(1-\beta)\pi_b,
\]
and relative ratio
\[
\rho_\beta=\frac{\rho}{1-\beta+\beta\rho},
\]
the relative Pearson divergence is
\[
\mathrm{PE}_\beta(\pi,\pi_b)=\mathbb E_{a\sim\pi_\beta}\left[\frac{1}{2}(\rho_\beta(a)-1)^2\right].
\]
Because \(\rho_\beta\in[0,1/\beta)\), the paper presents this as a bounded and numerically stable alternative to KL-based regularization in PPO-like algorithms [2010.03290].

A more recent discrete-energy formulation introduces **ratio divergence**
\[
L(\hat{P}, P;\theta) = \sum_{x',x \in X} \hat{P}(x')\, P(x;\theta) \left( \log \frac{\hat{P}(x')P(x;\theta)}{P(x';\theta)\hat{P}(x)} \right)^2,
\]
which compares target and model through pairwise probability ratios. For energy-based models this cancels partition functions and becomes alignment of model and target energy differences. The paper further proves the decomposition
\[
L(\hat{P}, P;\theta)= 2\,D_{\mathrm{KL}[\hat P\|P]\, D_{\mathrm{KL}[P\|\hat P] + D_{\mathrm{KL}_2}[\hat P\|P] + D_{\mathrm{KL}_2}[P\|\hat P],
\]
and derives a lower bound on expected Metropolis–Hastings acceptance in terms of \(L\) [2409.07679].

RADAR, by contrast, does not define a standalone metric called RD. Its local quantity most closely resembling an RD term is the relative distance descriptor
\[
d^{(l)}(x,x') = \frac{\|\mathbf v_{\mathrm{sep}}^{(l)}\|+\|\mathbf v_{\mathrm{detour}}^{(l)}\|-\|\mathbf v_{\mathrm{traj}}^{(l)}\|}{\max(\|\mathbf v_{\mathrm{traj}}^{(l)}\|,\epsilon)},
\]
while the global score is a weighted symmetric KL divergence between fitted distributions of trajectory descriptors [2605.23028].

These model-oriented formulations indicate a broader methodological shift: divergence is increasingly defined through the behavior of a learning system, not solely through an abstract geometry on distributions. A plausible implication is that in modern ML, “relative divergence” often denotes a task-conditioned diagnostic rather than a universal discrepancy measure.

## 7. Conceptual synthesis

The arXiv literature does not support a single encyclopedia-style formula for Relative Divergence. Instead, it supports a taxonomy. In information theory and quantum theory, RD is usually a divergence between states or distributions relative to a reference state, often with Rényi parameterization [1206.2459] [1308.5961]. In order-theoretic settings, RD is an increment-based comparison of grading functions on chains and posets, with MRDP as its variational principle [2510.04314]. In geometric group theory, RD is a pair invariant measuring avoidance geometry relative to a subgroup [1406.4232]. In machine learning, RD-like quantities may be defined through relative density ratios, model risks, representation trajectories, or target-energy differences [1106.4729] [2310.01109] [2409.07679].

What unifies these otherwise disparate objects is not a single formula but a structural theme: each construction compares one object to another through a relative reference mechanism that discards absolute scale and emphasizes constrained comparability. The reference may be a second density, a support projector, a null grading, a subgroup neighborhood, a hypothesis minimizing mixture risk, or a smoothed denominator. This suggests that “relative divergence” functions less as a fixed technical term than as a family of domain-specific comparison principles.

Source: https://www.emergentmind.com/topics/relative-divergence-rd