Papers
Topics
Authors
Recent
Search
2000 character limit reached

Computational Max-Divergence

Updated 12 July 2026
  • Computational max-divergence is a family of divergence measures that quantifies the maximal KL error in approximating a target distribution by a model, applicable in both classical and quantum contexts.
  • It leverages geometric formulations such as information geometry, Voronoi polytopes, and toric models to transform complex divergence maximization into tractable optimization problems.
  • Efficient measurement constraints in quantum settings lead to a computational formulation that bounds distinguishability and guides operational performance in state discrimination.

Computational max-divergence denotes a family of extremal divergence constructions rather than a single invariant. In statistical model approximation, it is the maximal information divergence from a model MΔ(X)M\subseteq \Delta(X), namely the largest Kullback–Leibler error incurred when approximating an arbitrary target distribution by MM. In quantum information, closely related notions appear as maximal quantum ff-divergences, defined through reverse tests and characterized as largest monotone quantum ff-divergences, and as computational max-divergences induced by families of efficiently implementable binary measurements rather than the full positive semidefinite cone (Montufar et al., 2013, Matsumoto, 2013, Hiai, 2018, Yángüez et al., 25 Sep 2025). This multiplicity of meanings is not merely terminological: each setting fixes a different optimization domain, projection geometry, and operational interpretation.

1. Basic quantities and conceptual scope

For a finite state space XX, the classical maximum-information-divergence problem starts from a model MΔ(X)M\subseteq \Delta(X) and the Kullback–Leibler divergence

D(pq)=xXp(x)logp(x)q(x).D(p\|q)=\sum_{x\in X} p(x)\log\frac{p(x)}{q(x)}.

The divergence from a distribution pp to the model is

D(pM):=infqMD(pq),D(p\|M):=\inf_{q\in M}D(p\|q),

and the quantity of interest is the worst-case approximation error

DM:=maxpΔ(X)D(pM).D_M:=\max_{p\in\Delta(X)}D(p\|M).

A distribution MM0 achieving the infimum is called the MM1-projection of MM2 onto MM3 (Montufar et al., 2013).

In this setting, MM4 measures the worst-case model misspecification error: one picks the distribution hardest for the model to approximate, computes the minimum KL divergence to the model, and maximizes over all targets. The same quantity admits a geometric reading. The model is regarded as a subset of the probability simplex, the MM5-projection is the KL-closest point in the model closure, and the maximal divergence is a global “radius” of the model in information geometry (Montufar et al., 2013).

The phrase “maximal” is used differently in quantum work. There, a maximal quantum MM6-divergence is defined either as the minimum classical MM7-divergence over all reverse tests, or by an operator formula when MM8 is operator convex, and it is maximal in the sense of being the largest quantum MM9-divergence compatible with data processing and agreement with the classical divergence on commuting states (Matsumoto, 2013). In the computationally constrained quantum setting, the ordering ff0 is replaced by a cone order induced by efficient binary measurements, so the “max” refers to distinguishability certified by efficient observers (Yángüez et al., 25 Sep 2025). This suggests that “computational max-divergence” is best understood as a class of extremal divergence problems indexed by the admissible model or test family.

2. Exponential families and optimization over ff1

For an exponential family ff2 on a finite set ff3, specified by a strictly positive reference measure ff4 and a sufficient-statistics matrix ff5, the ff6-projection ff7 of a probability measure ff8 is characterized by the moment-matching identity

ff9

It satisfies the Pythagorean identity

ff0

and therefore

ff1

where

ff2

Equivalently, ff3 maximizes ff4 among all probability measures satisfying ff5 (0912.4660).

Local maximizers of ff6 obey a strong projection property. If ff7 is a local maximizer with support ff8 and ff9 is its XX0-projection, then

XX1

Thus the maximizer is obtained by restricting its projection point to its support and renormalizing. There is also a support bound,

XX2

The paper further shows that XX3 can be paired with a disjoint-support measure XX4 satisfying XX5, sharing the same XX6-projection, with

XX7

A useful identity is

XX8

which couples the divergence values of the two disjoint-support maximizers (0912.4660).

A central computational reformulation replaces optimization over all probability measures by optimization over the kernel of the sufficient-statistics map. Every nonzero XX9 decomposes uniquely as MΔ(X)M\subseteq \Delta(X)0 with disjoint supports, and with normalization MΔ(X)M\subseteq \Delta(X)1, one introduces

MΔ(X)M\subseteq \Delta(X)2

Maximizing MΔ(X)M\subseteq \Delta(X)3 over MΔ(X)M\subseteq \Delta(X)4 is equivalent to maximizing MΔ(X)M\subseteq \Delta(X)5: global maximizers correspond bijectively, and local maximizers of MΔ(X)M\subseteq \Delta(X)6 yield local maximizers of MΔ(X)M\subseteq \Delta(X)7 (0912.4660).

When MΔ(X)M\subseteq \Delta(X)8 has integer entries, the first-order conditions become algebraic after fixing a sign vector MΔ(X)M\subseteq \Delta(X)9. The resulting quasi-critical equations exponentiate to binomial equations, placing the problem in the setting of toric and lattice-ideal computation. Two algorithms are proposed. The first enumerates sign vectors of D(pq)=xXp(x)logp(x)q(x).D(p\|q)=\sum_{x\in X} p(x)\log\frac{p(x)}{q(x)}.0, computes bases of D(pq)=xXp(x)logp(x)q(x).D(p\|q)=\sum_{x\in X} p(x)\log\frac{p(x)}{q(x)}.1, forms binomial systems, and solves them via saturation and primary decomposition; TOPCOM, Singular, solve.lib, and 4ti2 are explicitly noted. The second uses the projection-point parameterization

D(pq)=xXp(x)logp(x)q(x).D(p\|q)=\sum_{x\in X} p(x)\log\frac{p(x)}{q(x)}.2

and solves D(pq)=xXp(x)logp(x)q(x).D(p\|q)=\sum_{x\in X} p(x)\log\frac{p(x)}{q(x)}.3 for the parameters D(pq)=xXp(x)logp(x)q(x).D(p\|q)=\sum_{x\in X} p(x)\log\frac{p(x)}{q(x)}.4. The first method is easier when D(pq)=xXp(x)logp(x)q(x).D(p\|q)=\sum_{x\in X} p(x)\log\frac{p(x)}{q(x)}.5 is small, whereas the second is better suited when the dimension of the exponential family is small (0912.4660).

These methods resolve nontrivial examples. For the binary independence model on D(pq)=xXp(x)logp(x)q(x).D(p\|q)=\sum_{x\in X} p(x)\log\frac{p(x)}{q(x)}.6, the two global maximizers are

D(pq)=xXp(x)logp(x)q(x).D(p\|q)=\sum_{x\in X} p(x)\log\frac{p(x)}{q(x)}.7

For the binary D(pq)=xXp(x)logp(x)q(x).D(p\|q)=\sum_{x\in X} p(x)\log\frac{p(x)}{q(x)}.8-D(pq)=xXp(x)logp(x)q(x).D(p\|q)=\sum_{x\in X} p(x)\log\frac{p(x)}{q(x)}.9 model, the maximal value is

pp0

and for the independence model of cardinalities pp1, the maximal value is

pp2

with a unique global maximizer up to symmetry (0912.4660).

3. Linear and toric models: Voronoi geometry and chamber-complex algorithms

A later development treats maximum information divergence through logarithmic Voronoi polytopes. For a model pp3, the divergence from pp4 to the model is

pp5

and the maximum information divergence is

pp6

For fixed pp7, the map pp8 is strictly convex on the simplex (Alexandr et al., 2023).

For a pp9-dimensional linear model D(pM):=infqMD(pq),D(p\|M):=\inf_{q\in M}D(p\|q),0, the principal theorem states that the maximum divergence is always achieved at a vertex of the logarithmic Voronoi polytope D(pM):=infqMD(pq),D(p\|M):=\inf_{q\in M}D(p\|q),1, where D(pM):=infqMD(pq),D(p\|M):=\inf_{q\in M}D(p\|q),2 itself is a vertex of D(pM):=infqMD(pq),D(p\|M):=\inf_{q\in M}D(p\|q),3: D(pM):=infqMD(pq),D(p\|M):=\inf_{q\in M}D(p\|q),4 The proof uses strict convexity on each convex polytope D(pM):=infqMD(pq),D(p\|M):=\inf_{q\in M}D(p\|q),5, together with a co-circuit description of the vertices. If D(pM):=infqMD(pq),D(p\|M):=\inf_{q\in M}D(p\|q),6 is a normalized co-circuit, then the corresponding vertex D(pM):=infqMD(pq),D(p\|M):=\inf_{q\in M}D(p\|q),7 satisfies

D(pM):=infqMD(pq),D(p\|M):=\inf_{q\in M}D(p\|q),8

which is linear in D(pM):=infqMD(pq),D(p\|M):=\inf_{q\in M}D(p\|q),9. Hence the global problem reduces to maximizing a linear function over a polytope (Alexandr et al., 2023).

For toric models, including discrete exponential families, the logarithmic Voronoi polytope at DM:=maxpΔ(X)D(pM).D_M:=\max_{p\in\Delta(X)}D(p\|M).0 is the affine fiber

DM:=maxpΔ(X)D(pM).D_M:=\max_{p\in\Delta(X)}D(p\|M).1

A local maximizer of DM:=maxpΔ(X)D(pM).D_M:=\max_{p\in\Delta(X)}D(p\|M).2 must be a projection point, and every such projection point is a complementary vertex of DM:=maxpΔ(X)D(pM).D_M:=\max_{p\in\Delta(X)}D(p\|M).3, where DM:=maxpΔ(X)D(pM).D_M:=\max_{p\in\Delta(X)}D(p\|M).4 is the maximum likelihood estimate of the maximizer. The support of any maximizer satisfies

DM:=maxpΔ(X)D(pM).D_M:=\max_{p\in\Delta(X)}D(p\|M).5

This sparsity statement aligns with the support bound for exponential-family maximizers obtained by kernel methods (Alexandr et al., 2023).

The combinatorial types of the fibers DM:=maxpΔ(X)D(pM).D_M:=\max_{p\in\Delta(X)}D(p\|M).6 are organized by the chamber complex DM:=maxpΔ(X)D(pM).D_M:=\max_{p\in\Delta(X)}D(p\|M).7. For DM:=maxpΔ(X)D(pM).D_M:=\max_{p\in\Delta(X)}D(p\|M).8 in the relative interior of a chamber, the combinatorial type of DM:=maxpΔ(X)D(pM).D_M:=\max_{p\in\Delta(X)}D(p\|M).9 is constant, and the supports of vertices and faces are constant on that chamber. This yields an explicit algorithm: compute equations of the toric variety, compute the chamber complex, choose one representative MM00 per chamber, identify complementary vertex/face pairs, parametrize the corresponding line families, substitute into the equations of the toric variety, and solve the resulting systems using numerical algebraic geometry, with Bertini and HomotopyContinuation.jl given as examples (Alexandr et al., 2023).

The same framework produces explicit formulas in structured families. For twisted Veronese models MM01,

MM02

For box models MM03,

MM04

For the conditional independence model MM05,

MM06

The paper also emphasizes that reducible models admit decomposition results for Voronoi polytopes and divergence, but that compatibility of maximizers can fail, so naive additive upper bounds are not always attained (Alexandr et al., 2023).

4. Neural-network statistical models and worst-case approximation bounds

The review of maximal information divergence from neural-network statistical models considers the independence model MM07, naive Bayes mixtures MM08, restricted Boltzmann machines with MM09 visible and MM10 hidden units, deep belief networks, and comparison classes of exponential families such as partition models and hierarchical/log-linear models (Montufar et al., 2013).

The main technical strategy is geometric and model-theoretic. If a complicated model MM11 contains a tractable submodel MM12, then

MM13

This allows upper bounds for hard hidden-variable models to be obtained from simpler exponential submodels. A second ingredient is a decomposition theorem for mixtures supported on disjoint blocks: if MM14 partitions MM15 and MM16, then the MM17-projection onto the mixture MM18 decomposes blockwise as

MM19

This enables exact or near-exact reduction of union models to tractable block components (Montufar et al., 2013).

Several closed-form or nearly closed-form expressions are available. For the partition model MM20 with coarseness MM21,

MM22

Its maximizers have support intersecting each block in at most one point, and exactly one point in a block of maximal size. For an exponential family of dimension MM23,

MM24

and if equality holds, then MM25 must be a homogeneous partition model. For the MM26-ary independence model,

MM27

while in general

MM28

The maximizers in the MM29-ary case are uniform distributions on MM30-ary codes of size MM31 and minimum distance MM32 (Montufar et al., 2013).

For hidden-variable neural-network models, the same submodel method yields explicit worst-case bounds. For the naive Bayes model MM33, a general bound is

MM34

under the condition MM35. In the binary case,

MM36

For restricted Boltzmann machines with hidden state-space sizes MM37, if

MM38

then

MM39

and for binary units with MM40,

MM41

These formulas subsume the naive Bayes and independence cases as special cases at MM42 and MM43 (Montufar et al., 2013).

For deep belief networks with finite-valued units, the paper gives a new result for deep and narrow architectures. If the network has MM44 layers of width MM45, unit cardinalities MM46, and if

MM47

with MM48 chosen so that

MM49

then

MM50

In the binary case, if the network has

MM51

layers of size

MM52

then

MM53

The paper notes that the binary case follows prior work, whereas the non-binary case is new. Within this approximation-theoretic framework, universal approximation corresponds to vanishing maximal divergence, and deepening the network can improve approximation power, but narrow DBNs may still fail to be universal approximators (Montufar et al., 2013).

5. Maximal quantum MM54-divergence and reverse-test formulations

In finite dimensions, a reverse test of a pair of positive operators MM55 is a triple MM56 where MM57 is a positive trace-preserving map from classical probability measures to operators, MM58 are probability distributions on a finite set MM59, and

MM60

The maximal quantum MM61-divergence is defined by the optimization problem

MM62

Among all quantum divergences MM63 satisfying data processing under CPTP maps and agreement with the classical MM64-divergence on commuting states, MM65 is the largest one: any such MM66 obeys

MM67

The same paper gives the dual characterization

MM68

showing that MM69 is a pointwise supremum of linear functionals (Matsumoto, 2013).

When MM70 is proper, lower semicontinuous, operator convex, MM71, and MM72, a closed formula is available: MM73 If MM74, the correction term vanishes and the expression reduces to

MM75

The same work establishes joint convexity, positive homogeneity, direct-sum additivity, and lower semicontinuity. It also proves that for operator convex MM76, the maximum MM77-divergence of the outcome distributions of a measurement is strictly less than MM78, while this fails in general without operator convexity; the counterexample MM79 corresponds to total variation distance (Matsumoto, 2013).

In the von Neumann algebra setting, maximal MM80-divergence is developed for normal positive functionals using Haagerup’s MM81-space. If MM82, then there exists a unique operator MM83 such that

MM84

and for operator convex MM85,

MM86

For arbitrary MM87, the definition is extended by approximation. The central reverse-test theorem states

MM88

and the minimum is attained. The theory further yields a general integral expression, monotonicity under unital positive normal maps, joint convexity, joint lower semicontinuity in the norm topology, and martingale convergence

MM89

for increasing nets of von Neumann subalgebras (Hiai, 2018).

This operator-algebraic theory also clarifies the meaning of “maximal.” The maximal divergence dominates the standard one, coincides with it in the commuting case, and for MM90 yields the Belavkin–Staszewski relative entropy. In particular,

MM91

with equality when MM92 and MM93 commute (Hiai, 2018).

6. Efficient measurements, many-copy regularization, and algorithmic computation

A recent quantum-information formulation defines computational max-divergence relative to a family of polynomially generated sets of effect operators

MM94

where each MM95 consists of effects implementable by at most MM96 elementary gates for some polynomial MM97. The family is required to be closed under complements and tensor products, and generates a cone

MM98

which is proved to be a proper cone: convex, closed, pointed, and solid. The induced partial order is defined by MM99 (Yángüez et al., 25 Sep 2025).

Using the general ff00-max-divergence formalism, the computational max-divergence is defined for ff01 with ff02 and ff03 by

ff04

The equality with the dual expression follows from standard conic duality for proper cones. Operationally, this quantity measures the strongest bias obtainable in state discrimination when restricted to efficient binary tests. The paper also stresses that the converse to the ordinary positive-semidefinite implication need not hold: ff05 but not vice versa, so the computational order only captures positivity as seen through efficient measurements (Yángüez et al., 25 Sep 2025).

The same work defines the computational measured Rényi divergence over the set ff06 of efficient binary POVMs and proves the central equivalence

ff07

Thus the computational max-divergence is exactly the ff08 limit of the computational measured Rényi divergence, mirroring the unconstrained identity for unrestricted measurements. The paper also proves super-additivity under a strictly super-additive polynomial bound, approximate sub-additivity under finite preparation cost ff09, and the lower bound

ff10

where ff11 is the computational trace distance (Yángüez et al., 25 Sep 2025).

For many-copy regimes, regularized versions are introduced. The regularized computational max-divergence satisfies

ff12

with the equality of limit and supremum obtained from super-additivity and Fekete’s lemma. The measured regularizations obey the ordering

ff13

and a computational Stein bound gives the regularized computational measured relative entropy an operational role as an upper bound on the achievable type-II error exponent under efficient binary measurements (Yángüez et al., 25 Sep 2025).

The same framework induces resource measures. For a free-state set ff14,

ff15

The measured resource divergence is monotone under efficient free operations and faithful if ff16 is closed. An asymptotic continuity bound is proved when ff17 is closed, convex, bounded, and contains a full-rank state. For entanglement, the computational measured relative entropy of entanglement

ff18

acts as an entanglement monotone under efficient LOCC, and if the Bell projector ff19 is efficiently generated, then

ff20

The paper also constructs explicit separations using pseudoentangled states, showing that computational entanglement quantities can be negligible while the corresponding information-theoretic quantities remain maximal (Yángüez et al., 25 Sep 2025).

Related algorithmic work addresses the direct computation of maximal quantum ff21-divergence. Using the block-encoding + QSVT paradigm, one can compute

ff22

for operator-convex ff23, under purification access or sample access. The method constructs block-encodings of ff24 and ff25, forms

ff26

approximates functions such as ff27 either by direct QSVT polynomial approximation or by Stieltjes/resolvent representations, and then estimates the trace. For general operator-convex ff28, the Löwner–Kraus representation

ff29

reduces computation to resolvent kernels implemented with the same block-encoding and inversion primitives (Dinh et al., 13 Nov 2025).

Across these lines of work, computational max-divergence is a unifying extremal principle with several concrete realizations: worst-case KL approximation error for statistical models, reverse-test maximality for quantum ff30-divergences, and efficient-observer distinguishability under computational constraints. The common structure is an outer maximization over admissible targets or tests together with an inner projection, comparison, or conic-order minimization; the differences lie in which model class, geometry, and operational restrictions are built into that extremal problem.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Computational Max-Divergence.