---
title: Maximum Relative Divergence Principle
url: https://www.emergentmind.com/topics/maximum-relative-divergence-principle-mrdp
type: topic
---

# Maximum Relative Divergence Principle

The Maximum Relative Divergence Principle (MRDP) is a variational principle that selects extremal objects by optimizing a relative-entropy functional under explicit constraints. Across the literature, it appears in equivalent sign conventions: either as maximization of a “relative entropy” $S=-D$ or as minimization of a divergence such as Kullback–Leibler divergence or Umegaki relative entropy. In the finite discrete setting, MRDP asks for those distributions $p$ in the probability simplex that are most incompatible with a model $M$, as measured by $D(p\Vert M)$ [2308.15598]. In inference settings with a prior $q$, it chooses the posterior closest to $q$ subject to the new information, which is the standard maximum relative entropy or minimum relative entropy formulation [1708.00689, 1407.7766]. In many-party quantum systems, the same principle is instantiated by maximizing divergence from a hierarchical Gibbs family, thereby quantifying correlation content not captured by specified interaction scales [1406.0833].

## 1. General definition and sign conventions

MRDP is used in two mathematically equivalent forms. In the classical maximum relative entropy formulation, one maximizes
$$
S[p,q] = -\sum_x p(x)\log\!\left(\frac{p(x)}{q(x)}\right)
$$
subject to normalization and moment constraints, or equivalently minimizes $D_{KL}(p\Vert q)$ under the same constraints [1708.00689]. The corresponding Lagrangian yields the exponential-family solution
$$
p(x)=q(x)\exp\!\left(\lambda_0+\sum_{i=1}^m \lambda_i f_i(x)\right),
$$
with an analogous continuous formulation [1708.00689]. In this sense, “maximum relative divergence” and “maximum relative entropy” are synonymous because relative entropy is $D_{KL}(p\Vert q)$ up to sign [1708.00689].

A closely related formulation appears in quantum theory through the Umegaki relative entropy
$$
D(\rho\Vert\sigma)=\operatorname{tr}\big[\rho(\log\rho-\log\sigma)\big],
$$
defined when $\operatorname{ker}(\sigma)\subseteq \operatorname{ker}(\rho)$ and otherwise equal to $+\infty$ [1406.0833]. In the measurement-update setting, one writes either
$$
\sigma^*=\operatorname*{arg\,min}_{\sigma\in\mathcal{C}} D(\rho\Vert\sigma)
$$
or
$$
\sigma^*=\operatorname*{arg\,max}_{\sigma\in\mathcal{C}} S(\rho,\sigma),\qquad S(\rho,\sigma):=-D(\rho\Vert\sigma),
$$
and the two formulations are explicitly identified as equivalent [1407.7766].

A different, model-criticism-oriented form arises when a model $M$ is fixed and one defines
$$
D_M(p):=\min_{q\in M} D(p\Vert q),\qquad D(M):=\max_{p\in\Delta_{n-1}} D_M(p).
$$
Here MRDP identifies the data distributions that are maximally incompatible with the model [2308.15598]. This suggests that the principle has two complementary readings: an entropic-projection reading, where one finds the least informative update compatible with constraints, and a worst-case reading, where one finds the most non-model element relative to a model class.

## 2. Entropic projections, maximum entropy, and Bayesian updating

A central theme in MRDP is the relation between entropy maximization and divergence minimization. In the Bayesian-inference formulation, Giffin and Caticha show that Bayesian updating is a special case of MRDP: if $q(\theta,x)=q(\theta)q(x\mid\theta)$ is the joint prior and the observation $x=x^*$ is encoded as the hard data constraint $p(x)=\delta(x-x^*)$, then maximizing relative entropy yields
$$
p(\theta)\propto q(\theta)\,q(x^*\mid\theta),
$$
which is Bayes’ rule [1708.00689]. More general soft constraints lead to exponential-family posteriors of the form
$$
p(\theta)\propto q(\theta)\exp\!\left\{\sum_i \lambda_i f_i(\theta)\right\}
$$
[1708.00689].

The same geometry appears in Gibbs families of quantum states. For a finite-dimensional $C^*$-algebra $A$ and a real vector space $H\subset A_{sa}$ of self-adjoint matrices, the Gibbs family is
$$
\mathcal{E}:=R(H):=\left\{\rho=\frac{e^h}{\operatorname{tr}(e^h)}\,\middle|\,h\in H\right\},
$$
or, with $H=\operatorname{span}\{H_1,\dots,H_m\}$,
$$
\rho(\theta)=\frac{\exp\big(\sum_{i=1}^m \theta_i H_i\big)}{Z(\theta)},\qquad Z(\theta):=\operatorname{tr}\!\left(\exp\big(\sum_{i=1}^m \theta_i H_i\big)\right)
$$
[1406.0833]. For such families, the divergence from the model is
$$
D(\rho\Vert\mathcal{E})=\inf_{\sigma\in\mathcal{E}} D(\rho\Vert\sigma),
$$
and there is a unique information projection $\pi_{\mathcal{E}}(\rho)$ in the $rI$-closure characterized by moment matching,
$$
\langle h,\rho\rangle=\langle h,\pi_{\mathcal{E}}(\rho)\rangle\qquad \text{for all } h\in H
$$
[1406.0833].

The projection theorem states
$$
D(\rho\Vert\mathcal{E})=D(\rho\Vert\pi_{\mathcal{E}}(\rho))=\min_{\sigma\in(\mathcal{E})} D(\rho\Vert\sigma),
$$
and the Pythagorean theorem gives
$$
D(\rho\Vert\sigma)=D(\rho\Vert\pi_{\mathcal{E}}(\rho))+D(\pi_{\mathcal{E}}(\rho)\Vert\sigma),\qquad \forall \sigma\in(\mathcal{E})
$$
[1406.0833]. The associated maximum-entropy characterization is
$$
\pi_{\mathcal{E}}(\rho)=\arg\max\{H(\tau)\mid \langle h,\tau\rangle=\langle h,\rho\rangle,\ \forall h\in H\},
$$
so the entropy gap and the divergence coincide:
$$
D(\rho\Vert\mathcal{E})=H\big(\pi_{\mathcal{E}}(\rho)\big)-H(\rho)
$$
[1406.0833]. In this form, MRDP operationalizes irreducible structure as a distance from the maximum-entropy state that matches the prescribed moments.

## 3. Hierarchical models and many-party quantum correlations

One of the most developed instantiations of MRDP appears in the analysis of many-party quantum correlations. For $N$ units with local algebras $A_i\subset M_{n_i}$ and total algebra $A_{[N]}=A_1\otimes\cdots\otimes A_N$, the paper defines pure factor spaces $\tilde A_v$ and the decomposition
$$
A_{[N]}=\bigoplus_{v\subset[N]} \tilde A_v
$$
[1406.0833]. Given a downward-closed hypergraph $U\subset 2^{[N]}$ covering $[N]$, the hierarchical Hamiltonian subspace is
$$
\tilde U:=\bigoplus_{v\in U}\tilde A_v,
$$
and the corresponding hierarchical Gibbs family is
$$
\mathcal{E}_U:=R(\tilde U\cap A_{[N]})
$$
[1406.0833]. The canonical $k$-body hierarchy is obtained from
$$
U_k:=\bigcup_{\ell=0}^k \binom{[N]}{\ell},
$$
whose Gibbs family $\mathcal{E}_k$ coincides with the Gibbs family of $k$-local Hamiltonians [1406.0833].

In this setting, MRDP identifies the “most non-$\mathcal{M}$” correlations by maximizing $D(\rho\Vert\mathcal{M})$ over states $\rho$ [1406.0833]. The quantities
$$
c_k(\rho):=H(\pi_{\mathcal{E}_k}(\rho))-H(\rho)=D(\rho\Vert\mathcal{E}_k)
$$
measure all correlations in $\rho$ that cannot be observed in any $k$-party subsystem [1406.0833]. Irreducible $k$-party correlation is then defined by
$$
C_k(\rho):=c_{k-1}(\rho)-c_k(\rho),\qquad 2\le k\le N,
$$
with the divergence identity
$$
C_k(\rho)=D\big(\pi_{\mathcal{E}_k}(\rho)\Vert \pi_{\mathcal{E}_{k-1}}(\rho)\big)
$$
[1406.0833].

For the independence model,
$$
c_1(\rho)=D(\rho\Vert\mathcal{E}_1)=D\big(\rho\Vert \rho_1\otimes\cdots\otimes \rho_N\big)=\sum_{i=1}^N H(\rho_i)-H(\rho),
$$
so total correlation equals multi-information [1406.0833]. For $N=2$, mutual information is
$$
I(\rho)=H(\rho_1)+H(\rho_2)-H(\rho)=D\big(\rho\Vert \rho_1\otimes\rho_2\big)
$$
[1406.0833].

The paper emphasizes three classical-versus-quantum differences in hierarchical models: missing factorization, discontinuity, and reduction of uncertainty [1406.0833]. In the classical case, distributions with at most $k$-party interactions admit multiplicative factorization
$$
p(x)=\prod_{\nu\subset[N],\,|\nu|=k} \psi_\nu(x_\nu),
$$
whereas the quantum hierarchical model has no nontrivial multiplicative factorization beyond product structure and must be defined through exponential families over $\tilde U$ [1406.0833]. Classical divergence $D(\rho\Vert\mathcal{E}_k^{\mathrm{cl}})$ is continuous for all $k$, but in the quantum case $D(\rho\Vert\mathcal{E})$ can be discontinuous when the $rI$-closure is not norm-closed; a stated example is that $c_2$ for three qubits is discontinuous at the GHZ state [1406.0833]. Finally, local maximizers satisfy distinct size bounds:
$$
|\operatorname{supp}(p)|\le \dim_{\mathbb{R}}\mathcal{E}+1 \qquad \text{(classical)},
$$
$$
\operatorname{rk}(\rho)\le \sqrt{\dim_{\mathbb{R}}\mathcal{E}+1} \qquad \text{(quantum)},
$$
which the paper interprets as a reduction of uncertainty for quantum maximizers [1406.0833].

A concrete global-maximizer result is given for separable two-qubit states:
$$
I(\rho)\le \log 2,
$$
with equality if and only if $\rho$ is local-unitary equivalent to
$$
\frac12\big(|0\rangle\langle 0|\otimes |0\rangle\langle 0|+|1\rangle\langle 1|\otimes |1\rangle\langle 1|\big)
$$
[1406.0833]. The paper states that all separable maximizers are classically correlated Bell-diagonal mixtures of two Bell states [1406.0833].

## 4. Linear models, toric models, and logarithmic Voronoi geometry

MRDP has also been formulated as a geometric optimization problem over linear and toric models. For a model $M\subset \Delta_{n-1}$, one studies
$$
D_M(p):=\min_{q\in M} D(p\Vert q),\qquad D(M):=\max_{p\in\Delta_{n-1}} D_M(p)
$$
[2308.15598]. For toric models, the unique minimizer is the maximum-likelihood estimate, and Birch’s Theorem states that it is the unique point solving $Ap=Aq$ inside the model [2308.15598].

A key geometric object is the logarithmic Voronoi polytope
$$
Q_q=\{u\in \Delta_{n-1}: Au=Aq\},
$$
which is the cell of points whose MLE is $q$ [2308.15598]. The paper proves that for linear or toric models, the maximum of $D_M(u)$ restricted to $Q_q$ is achieved at a vertex of $Q_q$, because $D(u\Vert q)$ is strictly convex in $u$ [2308.15598]. For linear models, this combines with co-circuit geometry to yield a boundary-attainment theorem: the maximum divergence is achieved at a vertex of a logarithmic Voronoi polytope $Q_q$ with $q$ itself a vertex of the model [2308.15598]. The same paper states the support bound
$$
|\operatorname{supp}(p)|\le d=\dim(M_A)+1
$$
for toric models with $\operatorname{rank}(A)=d$ [2308.15598].

The toric case is more involved. The chamber complex $C_A$ partitions $\operatorname{conv}(A)$ into regions where the logarithmic Voronoi polytopes have identical combinatorial type [2308.15598]. Within a chamber, vertices of $Q_b$ are in bijection with certain linearly independent subsets $\sigma\subset[n]$, and MRDP candidates are characterized as projection points or complementary vertices satisfying geometric intersection conditions [2308.15598]. The paper’s algorithm combines chamber-complex combinatorics with numerical algebraic geometry: compute equations of the toric variety, compute the chamber complex, enumerate complementary vertex/face pairs, intersect parameterized line families with the toric variety, impose positivity, and maximize $D_{M_A}$ over the resulting semi-algebraic set [2308.15598].

Several exact values are given. For the binomial model of size $3$, the global MRDP maximizer is $(1/2,0,0,1/2)$ with $D=2\log 2$ [2308.15598]. For the independence model $2\times 2$, the only two projection points are $(1/2,0,0,1/2)$ and $(0,1/2,1/2,0)$, both achieving $D=\log 2$ [2308.15598]. For the conditional-independence model $M_{[12][23],\vec d}$,
$$
D(M_{[12][23],\vec d})=\min(\log d_1,\log d_3)
$$
[2308.15598]. For reducible hierarchical models, divergence obeys additive lower bounds and, under compatibility conditions, exact additivity across components [2308.15598].

This line of work presents MRDP as a method of model criticism, misspecification detection, and stress-testing [2308.15598]. A plausible implication is that the “most incompatible” data against a model are often sparse, because the maximizers concentrate on supports of small cardinality and frequently lie on the boundary of the simplex.

## 5. Grading functions on power sets, chain bundles, and posets

A distinct generalization of MRDP replaces probability distributions by grading functions on ordered structures. On a finite event space $\Omega$ with power set $W=2^\Omega$, a grading function is a set-monotone map $F$ satisfying
$$
w\subseteq v \Rightarrow F(w)\le F(v)
$$
[2207.07099]. On a maximal chain $MC=\{w_0,\dots,w_n\}$, the increments are $\Delta_iF=F(w_i)-F(w_{i-1})>0$, and the normalized grading function is
$$
\hat F(w):=\frac{F(w)-m_F}{\Delta_W F}
$$
with $\hat F(w_0)=0$ and $\hat F(w_n)=1$ [2207.07099].

For totally ordered chains, relative divergence is defined by
$$
D(F\Vert G)\big|_W:=\sum_k \Delta_kF \ln\!\left(\frac{\Delta_k G}{\Delta_k F}\right),
$$
which equals the negative KL divergence of the increment distributions [2207.07099]. If $G$ is the ordinal grading with unit increments, then
$$
D(F\Vert I)\big|_W=-\sum_k \Delta_kF \ln \Delta_kF,
$$
which is Shannon entropy after normalization [2207.07099]. On power sets, the paper defines
$$
D(F\Vert G)\big|_W:=\min_{MC\subset W} D(F\Vert G)\big|_{MC}
$$
and, with $G=N$ the cardinality grading, obtains
$$
D(F\Vert N)\big|_W=H(F)\big|_W+\Delta_W F \ln(\Delta_W F)
$$
[2207.07099].

MRDP on power sets selects admissible grading functions maximizing $D(F\Vert N)\big|_W$, which reduces in equilateral normalized cases to maximizing Shannon entropy [2207.07099]. For element-additive gradings
$$
F(w)=\sum_{x\in w} f(x),
$$
linear constraints yield the Lagrangian solution
$$
f(x)=e^{-1}e^{-\lambda^T \hat a_x},
$$
with multipliers determined by the nonlinear system
$$
A[e^{-\lambda^T \hat a_x}]_{x\in\Omega}=eM
$$
[2207.07099]. For cardinality-dependent gradings with fixed values at selected cardinalities, the solution is piecewise linear:
$$
F(w)=a_k+b_k|w|,\qquad |w|\in I_k,
$$
where
$$
b_k=\frac{M_k-M_{k-1}}{n_k-n_{k-1}}
$$
[2207.07099]. Under partition quotas, MRDP yields uniform allocation within each block:
$$
f(x)=\frac{M_k}{|X_k|},\qquad x\in X_k
$$
[2207.07099].

The direct-product analogue is developed for chain bundles $W=X_1\times\cdots\times X_R$ under the product order [2303.14261]. Relative divergence on the bundle is defined by
$$
D(F\mid G)\big|_W:=\min_{MC} D(F\mid G)\big|_{MC},
$$
and if $F$ and $G$ are additively separable then
$$
D(F\mid G)\big|_W=D(F_u\mid G_u)\big|_U + D(F_v\mid G_v)\big|_V
$$
for a decomposition $W=U\times V$ [2303.14261]. Height-dependent grading functions satisfy
$$
D(F\mid N)\big|_W=-\sum_{k=1}^K f(k)\ln f(k),
$$
so the unconstrained optimizer is linear in height:
$$
F(i)=m+\frac{M-m}{K}N(i)
$$
[2303.14261].

A further generalization to partially ordered sets introduces conjoined posets, serial and parallel block structures, and an “Insufficient Reason Principle with prior information” in which MRDP chooses the least-presuming grading function relative to a null grading function $G$ [2510.04314]. In this framework, RD is block-additive over serial composition and given by an infimum over maximal chains for even-sided split-chains [2510.04314]. The paper states that classic probability identities such as conditional probability, the independence product rule, the law of total probability, and Bayes’ theorem can be presented as MRDP solutions on conjoined posets [2510.04314]. Because this source is dated 2025-10-05, which is later than the present date, it should be treated cautiously as a reported extension rather than as settled background.

## 6. Applications and domain-specific interpretations

In Bayesian network learning, MRDP is used to assess whether a scoring rule updates away from the prior only when the data force it. For discrete Bayesian networks with Dirichlet hyperparameters $\alpha_{ijk}$, the local marginal likelihood is
$$
\text{BD}(X_i,\Pi_i)=\prod_{j=1}^{q_i}
\frac{\Gamma(\alpha_{ij})}{\Gamma(\alpha_{ij}+N_{ij})}
\prod_{k=1}^{r_i}
\frac{\Gamma(\alpha_{ijk}+N_{ijk})}{\Gamma(\alpha_{ijk})}
$$
[1708.00689]. BDeu chooses
$$
\alpha_{ijk}=\frac{\alpha}{r_i q_i},\qquad \alpha_{ij}=\frac{\alpha}{q_i},
$$
while BDs concentrates prior mass on observed parent configurations through
$$
\alpha_{ijk}=
\begin{cases}
\dfrac{\alpha}{r_i\tilde q_i}, & N_{ij}>0,\\[4pt]
0, & N_{ij}=0.
\end{cases}
$$
[1708.00689]. The paper argues that BDeu violates the maximum relative entropy principle in sparse data because its effective imaginary sample size changes across structures, whereas BDs preserves the imaginary sample size across structures, eliminates contributions from unobserved configurations, reduces $\alpha$-sensitivity, and is asymptotically score-equivalent to BDeu [1708.00689].

In quantum measurement theory, MRDP provides an information-theoretic characterization of the von Neumann–Lüders collapse rules. For a sharp observable $O=\sum_i \lambda_i P_i$, the weak measurement constraint is
$$
\mathcal{C}_{\mathrm{weak}}=\{\sigma\in\mathcal{D}\mid [P_i,\sigma]=0\ \text{for all } i\},
$$
and the unique minimizer of $D(\rho\Vert \sigma)$ on this set is
$$
\rho'=\sum_i P_i \rho P_i
$$
[1407.7766]. With fixed outcome probabilities $p_i$, the weighted strong constraint set yields the weighted Lüders rule
$$
\rho'_{\{p_i\}}=\sum_i p_i\,\frac{P_i\rho P_i}{\operatorname{Tr}(P_i\rho)},
$$
and the strong rule is recovered as
$$
\rho'_k=\frac{P_k\rho P_k}{\operatorname{Tr}(P_k\rho)}
$$
[1407.7766]. In the commuting case, the paper states that the quantum formulation reproduces classical MaxRelEnt updating, including Jeffrey’s rule [1407.7766].

In incompressible fluid mechanics, a principle of maximum entropy is proposed on the Hilbert space $H$ of $L^2$ divergence-free velocity fields on a periodic cube [2402.14240]. The relative entropy is
$$
S(\mu\Vert \nu)=-\int_H \log\!\left(\frac{d\mu}{d\nu}\right)\,d\mu,
$$
with $S(\mu\Vert \nu)=-\infty$ if $\mu$ is not absolutely continuous with respect to $\nu$ [2402.14240]. For a fixed-time energy–enstrophy surface
$$
G_e(t)=\{u\in X_t: |u|^2=e(t)\},
$$
the admissible measures satisfy $\mu_{e,t}\ll \nu_{e,t}$ and normalization, and the Euler–Lagrange condition yields
$$
f_{e,t}(u)=\frac{d\mu_{e,t}}{d\nu_{e,t}}=1.
$$
Thus the reference physical measure $\nu_{e,t}$ is the unique maximizer of $S(\mu\Vert \nu_{e,t})$ on the constrained set [2402.14240]. The paper connects this to stationary statistical solutions of the Navier–Stokes equations and to Kolmogorov’s “final statistics” for fully developed turbulence [2402.14240].

Operations Research applications appear in the grading-function literature. On power sets, MRDP is used for resource distribution with quotas, costs within blocks, and group testing; on direct products of chains it is applied to group service in queueing theory and resource distribution under constraints [2207.07099, 2303.14261]. The reported solutions are uniform under pure quota constraints, exponential-family under linear cost constraints, and piecewise linear under pinned-value or cardinality constraints [2207.07099, 2303.14261].

## 7. Structural differences, limitations, and open questions

Across these formulations, MRDP is not a single algorithm but a family of variational principles whose precise content depends on the choice of divergence, the ambient space, and the admissible constraint class. In the quantum hierarchical setting, the paper explicitly lists missing factorization, discontinuity, and reduction of uncertainty as the salient differences between quantum states and classical probability vectors [1406.0833]. In linear and toric models, the main technical limitation is computational: chamber complexes can be enormous, and the algebraic step requires solving polynomial systems with positivity constraints [2308.15598]. In Bayesian network scoring, the difficulty is prior sensitivity under sparsity, which is the basis for the critique of BDeu [1708.00689]. In Navier–Stokes, the central open modeling problem is the choice of the reference physical measure $\nu_{e,t}$ on the energy–enstrophy surface [2402.14240].

Several open questions are stated explicitly. For quantum hierarchical models, these include full characterization of global maximizers of $D(\rho\Vert \mathcal{E}_k)$ in the presence of entanglement, structural description of closures $(\mathcal{E}_k)$ and continuity regimes, and algorithmic reliability near non-maximal-rank or zero-temperature limits [1406.0833]. For toric models, higher-dimensional hierarchical families remain challenging, some model families remain conjectural, and compatibility constraints in reducible models can prevent simultaneous attainment of componentwise maxima [2308.15598]. For Navier–Stokes, the existence of self-similar homogeneous statistical solutions and a canonical infinite-dimensional construction of $\nu_{e,t}$ remain open [2402.14240]. In the poset-based grading-function program, a reported direction is a unified axiomatization of RD on arbitrary posets beyond block-additivity and split-infimum [2510.04314].

Taken together, these developments show that MRDP functions as a unifying principle for least-presumptive updating, extremal model criticism, and structured entropy maximization. In one direction it yields entropic projections compatible with moment, measurement, or support constraints; in the other it identifies maximally incompatible states or distributions relative to hierarchical, linear, or toric models. The recurring mathematical motifs are exponential-family structure, convexity or strict concavity, projection theorems, and boundary concentration of extremizers [1406.0833, 2308.15598].

Source: https://www.emergentmind.com/topics/maximum-relative-divergence-principle-mrdp