Relative Divergence: A Cross-Disciplinary Overview
- Relative Divergence is a family of metrics that compare objects relative to a reference structure, tailored to the ambient domain.
- Key formulations include KL and Rényi divergences, density-ratio estimations, grading functions on posets, and geometric invariants.
- Applications span information theory, quantum mechanics, ordered structures, geometric group theory, and model-oriented machine learning.
Searching arXiv for recent and foundational papers on “Relative Divergence” across the senses represented in the source material. I’ll synthesize the term as a cross-disciplinary concept, since the source material shows that “Relative Divergence (RD)” is used in multiple non-equivalent ways rather than as a single standardized definition. “Relative Divergence” (RD) is not a single universally standardized object across the arXiv literature. The term is used for several non-equivalent constructions that compare one object to another relative to a reference structure: probability distributions relative to a baseline distribution, quantum states relative to a second state, grading functions relative to another grading on an ordered set, finitely generated groups relative to a subgroup, and model-dependent discrepancies defined through a learning procedure. This multiplicity of usage suggests that RD is best understood as a family resemblance term: a divergence-like quantity whose precise meaning is fixed by the ambient category and by the operational question under study (Erven et al., 2012, Dukhovny, 5 Oct 2025, Tran, 2014).
1. Terminological scope and recurrent structural pattern
Across the literature, RD typically compares two objects by measuring how one departs from another after fixing a reference geometry, order, support condition, or hypothesis class. In some papers the object is a standard divergence between distributions; in others it is a quadratic discrepancy, a support-sensitive Rényi-type quantity, or a large-scale geometric invariant. The same abbreviation also collides with unrelated usages such as “property RD” for rapid decay, which is a representation-theoretic notion rather than a divergence (Boyer, 2013).
| Domain | Object compared | Representative construction |
|---|---|---|
| Information theory and statistics | relative to | Rényi divergence, KL divergence, relative Pearson divergence, relative extropy |
| Quantum information | relative to | Petz, sandwiched, maximal, minimal Rényi-type divergences |
| Ordered structures | grading function relative to | increment-based RD on chains and posets |
| Geometric group theory | group relative to subgroup | upper and lower relative divergence |
| Model-oriented ML | dataset/distribution relative to through a model | R-divergence, RDR, RADAR components |
A recurring pattern is that the divergence is not merely a pointwise comparison. It is often induced by an auxiliary structure: a likelihood ratio, a support projector, a partial order, a subgroup neighborhood, a target energy, or a hypothesis minimizing empirical risk. This structural dependence is explicit in the grading-function, group-theoretic, and model-oriented formulations (Dukhovny, 5 Oct 2025, Tran, 2014, Zhao et al., 2023).
2. Relative divergence in probability, information theory, and density-ratio methods
In the classical information-theoretic sense, RD is frequently identified either with Kullback–Leibler divergence, also called relative entropy, or with the broader Rényi divergence family. Rényi divergence of order 0 is defined by
1
with continuous extensions at 2. The order-3 case is the KL divergence
4
and the family is nondecreasing in 5, satisfies data processing, and interpolates between support-sensitive and worst-case notions of distinguishability (Erven et al., 2012). The same family is presented as a broad notion of “relative divergence” in another review, where 6, 7, and 8 is the logarithm of the essential supremum of 9 (Erven et al., 2010).
A separate statistical line develops relative divergences through density-ratio smoothing. Relative density-ratio estimation replaces the ordinary ratio
0
by the 1-relative ratio
2
where the denominator is the 3-mixture density 4. The associated relative Pearson divergence is
5
For 6, 7, so the target ratio is uniformly bounded; this is the paper’s main robustness argument. The resulting RuLSIF estimator uses a least-squares objective and yields a closed-form solution 8 for kernel models (Yamada et al., 2011).
Related relative constructions appear in more recent generative-model evaluation. There the relative density ratio is
9
which has bounded image 0. The paper couples this functional object to the scalar score
1
and proves that one-to-one transforms of the ordinary density ratio preserve 2-divergence under pushforward, so the RDR acts as a divergence-preserving one-dimensional summary of distributional discrepancy (Xu et al., 29 Oct 2025).
Another extropy-based branch defines relative extropy by the quadratic expression
3
together with directed extropy divergences 4 and 5 satisfying
6
This is positioned as an extropy-side analogue of the relation between inaccuracy and KL divergence, and is extended to residual and past lifetime settings through conditional densities (P. et al., 10 Mar 2025).
These formulations share a common design choice: direct comparison to 7 is replaced or supplemented by comparison to a smoothed, transformed, or structurally induced reference. This suggests that “relative” in RD often indicates regularization of the denominator or embedding of the comparison into a more stable functional setting.
3. Quantum relative divergences and Rényi-type endpoint structure
In quantum information, RD usually refers to a Rényi-type divergence between positive semidefinite operators or density matrices. A central distinction is between the traditional or Petz-type quantity
8
and the sandwiched quantum Rényi divergence
9
For 0, the sandwiched divergence satisfies the data-processing inequality and recovers several operationally important quantities: 1 as 2, 3 at 4, and 5 as 6 (Datta et al., 2013).
A key endpoint result concerns the 7-relative Rényi entropy
8
Datta and Leditzky proved that
9
holds when 0, but can fail if 1 is strict. Their conclusion is that the sandwiched divergence is not by itself a universal parent quantity for all operationally relevant relative entropies: one needs the sandwiched family for 2 and the traditional Rényi relative entropy for 3 (Datta et al., 2013).
Another quantum line studies sufficiency of Rényi divergences for state-pair equivalence. Classically, equality of 4 on an open interval determines interconvertibility of dichotomies. Quantum mechanically, known Rényi families are invariant under anti-unitary transformations, so no such family can be sufficient for CPTP-interconvertibility. The paper shows that the Petz and maximal quantum Rényi divergences remain insufficient even if one enlarges the notion of convertibility to positive trace-preserving maps, while giving evidence that the minimal or sandwiched family may be sufficient in that enlarged sense (Galke et al., 2023).
There is also a geometric equivalence result between relative 5-entropy and Rényi divergence. Under the escort transformation
6
the paper proves
7
and shows that projection theorems, Pythagorean inequalities, associated statistical families, and divergence-induced Riemannian metrics become equivalent under this correspondence (Karthik et al., 2017).
4. Relative divergence on ordered sets and grading functions
A distinct meaning of RD appears in the theory of grading functions on chains and posets. If 8 is a totally ordered chain and 9 are grading functions with increments
0
then the relative divergence of 1 from 2 on 3 is
4
When 5 is the index function with unit increments, this reduces to
6
so Shannon entropy appears as a special case of RD (Dukhovny, 5 Oct 2025).
The same literature treats RD as inherently dependent on order structure rather than as a universal distributional divergence. On posets, the definition is assembled from chainwise divergences using structural rules such as block-additivity on block-chains,
7
and chain-infinum on even-sided split-chains,
8
The paper explicitly states that there may be no universal RD definition for arbitrary posets; the definition must reflect the poset structure and the interdependence of maximal chains (Dukhovny, 5 Oct 2025).
This framework supports the Maximum Relative Divergence Principle (MRDP): among admissible grading functions with a given null grading 9, choose the one maximizing RD from 0. On finite chains with interpolation constraints, the maximizing grading is piecewise linear: 1 with
2
MRDP is then used to recover classical formulas such as
3
and
4
from optimization on conjoined event posets (Dukhovny, 5 Oct 2025).
An earlier paper develops the same program specifically for direct products of chains. There RD on a chain bundle is defined by taking the minimum over maximal chains,
5
with the natural grading 6. For additively separable grading functions, RD decomposes as a sum of componentwise terms: 7 This makes MRDP a direct generalization of the maximum-entropy principle to ordered multidimensional state spaces (Dukhovny, 2023).
5. Relative divergence in geometric group theory
In geometric group theory, “relative divergence” has a different meaning again. For a geodesic metric space 8 and subspace 9, one defines upper and lower relative divergence by measuring the difficulty of connecting points while avoiding neighborhoods of 0. With
1
and 2 the induced length metric on 3, the upper relative divergence is the family
4
taken over 5 satisfying 6 and 7. The lower relative divergence is
8
over points on 9 that are far apart in ambient distance but connected outside 0 (Tran, 2014).
For a finitely generated group 1 and subgroup 2, these become 3 and 4, quasi-isometry invariants of the pair 5. Upper RD generalizes Gersten’s divergence, and lower RD generalizes the lower divergence of Cooper–Mihalik. A major contrast emphasized in the paper is that classical lower divergence of a one-ended finitely generated group is only linear or exponential, whereas relative lower divergence can be any polynomial degree or exponential (Tran, 2014).
The paper derives comparison theorems. For a finitely generated normal subgroup 6 with one-ended quotient,
7
where 8 is upper distortion. For cyclic subgroups 9, lower RD is controlled by the divergence of the subgroup axis and the subgroup distortion: 00 The theory is then applied to CAT(0) groups and relatively hyperbolic groups, showing, for example, that for any polynomial or exponential function 01, there exists a CAT(0) pair 02 with 03 such that 04 (Tran, 2014).
This version of RD is not a divergence between measures at all. It is a large-scale geometric invariant measuring how the complement of a subgroup neighborhood stretches. Its inclusion under the same label illustrates how broad the term “relative divergence” has become across fields.
6. Model-oriented and recent machine-learning variants
Recent machine-learning work uses “divergence” language in explicitly model-dependent ways. One example is R-divergence, a model-oriented discrepancy between distributions 05 and 06, defined using the hypothesis 07 that minimizes risk on the mixture 08: 09 Its empirical estimator trains on pooled data 10 and compares empirical risks on the two parts. The paper emphasizes that this is not a universal 11-divergence; it depends on the hypothesis class, loss, and target function, so it measures whether two datasets are effectively the same for a specific learning model (Zhao et al., 2023).
Another use appears in reinforcement learning, where relative Pearson divergence regularizes policy updates. With importance ratio
12
mixture policy
13
and relative ratio
14
the relative Pearson divergence is
15
Because 16, the paper presents this as a bounded and numerically stable alternative to KL-based regularization in PPO-like algorithms (Kobayashi, 2020).
A more recent discrete-energy formulation introduces ratio divergence
17
which compares target and model through pairwise probability ratios. For energy-based models this cancels partition functions and becomes alignment of model and target energy differences. The paper further proves the decomposition
18
and derives a lower bound on expected Metropolis–Hastings acceptance in terms of 19 (Ishida et al., 2024).
RADAR, by contrast, does not define a standalone metric called RD. Its local quantity most closely resembling an RD term is the relative distance descriptor
20
while the global score is a weighted symmetric KL divergence between fitted distributions of trajectory descriptors (Cadet et al., 21 May 2026).
These model-oriented formulations indicate a broader methodological shift: divergence is increasingly defined through the behavior of a learning system, not solely through an abstract geometry on distributions. A plausible implication is that in modern ML, “relative divergence” often denotes a task-conditioned diagnostic rather than a universal discrepancy measure.
7. Conceptual synthesis
The arXiv literature does not support a single encyclopedia-style formula for Relative Divergence. Instead, it supports a taxonomy. In information theory and quantum theory, RD is usually a divergence between states or distributions relative to a reference state, often with Rényi parameterization (Erven et al., 2012, Datta et al., 2013). In order-theoretic settings, RD is an increment-based comparison of grading functions on chains and posets, with MRDP as its variational principle (Dukhovny, 5 Oct 2025). In geometric group theory, RD is a pair invariant measuring avoidance geometry relative to a subgroup (Tran, 2014). In machine learning, RD-like quantities may be defined through relative density ratios, model risks, representation trajectories, or target-energy differences (Yamada et al., 2011, Zhao et al., 2023, Ishida et al., 2024).
What unifies these otherwise disparate objects is not a single formula but a structural theme: each construction compares one object to another through a relative reference mechanism that discards absolute scale and emphasizes constrained comparability. The reference may be a second density, a support projector, a null grading, a subgroup neighborhood, a hypothesis minimizing mixture risk, or a smoothed denominator. This suggests that “relative divergence” functions less as a fixed technical term than as a family of domain-specific comparison principles.