Papers
Topics
Authors
Recent
Search
2000 character limit reached

Reverse-Pinsker-Type Bound

Updated 10 June 2026
  • Reverse-Pinsker-Type Bound is a framework that provides sharp upper bounds on f-divergences using total variation distance and pointwise constraints on the likelihood ratio.
  • The approach is applied to various divergences, including KL, Rényi, and χ², through convex functions, optimal three-point discrete models, and tight finite-support constructions.
  • These bounds have practical significance in statistical inference, large deviation theory, and privacy, linking divergence metrics with TV for enhanced analytical and operational insights.

A reverse-Pinsker-type bound refers to a family of sharp upper bounds on ff-divergences—in particular, those generated by convex functions such as the Kullback–Leibler (KL), Rényi, or χ² divergences—in terms of the total variation (TV) distance between two probability measures, sometimes incorporating pointwise constraints on the Radon–Nikodým derivative. While the classical Pinsker inequality gives a lower bound on TV in terms of divergence ("direct Pinsker"), its reverse addresses the fundamental question of how large a divergence can be if only the TV distance and certain extremal values of likelihood ratio are specified. These inequalities are operationally significant in information theory, statistics, large deviations, and local information geometry.

1. Core Definitions and Framework

Let PQP \ll Q denote two probability measures on a measurable space, and let f:[0,)(,]f : [0, \infty) \to (-\infty, \infty] be a convex function with f(1)=0f(1) = 0. The ff-divergence is

Df(PQ)=EQ[f(dPdQ)]D_f(P\|Q) = \mathbb{E}_Q\left[f\left(\frac{dP}{dQ}\right)\right]

and the total variation distance is

TV(P,Q)=12dP/dQ1dQ.\mathrm{TV}(P, Q) = \frac{1}{2} \int |dP/dQ - 1| \, dQ.

For pointwise control, let m=essinfQdP/dQm = \mathrm{ess\,inf}_Q \,dP/dQ and M=esssupQdP/dQM = \mathrm{ess\,sup}_Q \,dP/dQ (alternatively, the essential extrema of the "relative information" i(x)=logdPdQ(x)i(x) = \log \frac{dP}{dQ}(x)).

General Reverse-Pinsker-Type Bound (Binette (Binette, 2018)): For PQP \ll Q0 and PQP \ll Q1,

PQP \ll Q2

with equality achieved (optimality) for a discrete three-point model. This principle extends to many divergence types by choosing the appropriate convex generator PQP \ll Q3.

2. Major Classes of Reverse-Pinsker Inequalities

Reverse-Pinsker-type bounds have been developed for a variety of divergences and settings:

  • KL Divergence: For PQP \ll Q4, the sharp reverse-Pinsker bound is

PQP \ll Q5

where PQP \ll Q6, PQP \ll Q7 are the essential bounds on the Radon–Nikodým derivative reciprocal (Binette, 2018).

  • Rényi Divergence: For order PQP \ll Q8, the optimal bound is

PQP \ll Q9

where f:[0,)(,]f : [0, \infty) \to (-\infty, \infty]0, f:[0,)(,]f : [0, \infty) \to (-\infty, \infty]1, f:[0,)(,]f : [0, \infty) \to (-\infty, \infty]2, and

f:[0,)(,]f : [0, \infty) \to (-\infty, \infty]3

(Grosse et al., 20 Jan 2025). This is tight for TV distance f:[0,)(,]f : [0, \infty) \to (-\infty, \infty]4.

  • f:[0,)(,]f : [0, \infty) \to (-\infty, \infty]5-divergence: For f:[0,)(,]f : [0, \infty) \to (-\infty, \infty]6,

f:[0,)(,]f : [0, \infty) \to (-\infty, \infty]7

(Binette, 2018).

  • Hellinger and General f:[0,)(,]f : [0, \infty) \to (-\infty, \infty]8-divergences: Similar structure, with explicit formulas for Hellinger of order f:[0,)(,]f : [0, \infty) \to (-\infty, \infty]9 and other divergences, all subsumed under the Binette framework (Binette, 2018, Grosse et al., 20 Jan 2025).

A summary of specializations is shown below:

Divergence Generator f(1)=0f(1) = 00 Explicit Bound
KL f(1)=0f(1) = 01 See above
f(1)=0f(1) = 02 f(1)=0f(1) = 03 f(1)=0f(1) = 04
Rényi-f(1)=0f(1) = 05 f(1)=0f(1) = 06 f(1)=0f(1) = 07
Hellinger f(1)=0f(1) = 08 Explicit in f(1)=0f(1) = 09

3. Methodological Foundations and Proof Techniques

The foundational method involves a variational maximization of the divergence under TV and support constraints, usually reducing the extremal case to finite support models. The proofs typically follow these steps:

  1. Decomposition: Partition the space into ff0 and its complement.
  2. Convexity Application: Use convexity of ff1 and properties such as Jensen’s inequality, often combined with tight supporting lines at ff2.
  3. Explicit Model Construction: Build discrete measures (typically on two or three points) achieving the bound, confirming tightness.
  4. Operator methods: In the quantum or functional-analytic setting, operator-convexity replaces standard convexity (Rastegin, 2011).
  5. Alternative Proofs: For Rényi, ff3 (hockey-stick) representations integrate chordal approximations of ff4 weighted by the second derivative of ff5 (Grosse et al., 20 Jan 2025).

For the local (small perturbation) regime, Taylor expansion of the generator ff6 up to order three yields

ff7

for ff8, with ff9 a quadratic energy and Df(PQ)=EQ[f(dPdQ)]D_f(P\|Q) = \mathbb{E}_Q\left[f\left(\frac{dP}{dQ}\right)\right]0 controlling the third derivative, so

Df(PQ)=EQ[f(dPdQ)]D_f(P\|Q) = \mathbb{E}_Q\left[f\left(\frac{dP}{dQ}\right)\right]1

(Yu et al., 2024).

4. Extensions, Optimality, and Range

Reverse-Pinsker-type inequalities are sharp:

  • Optimality: For fixed TV and extremal ratios (Df(PQ)=EQ[f(dPdQ)]D_f(P\|Q) = \mathbb{E}_Q\left[f\left(\frac{dP}{dQ}\right)\right]2), the discrete three-point construction achieves equality (Binette, 2018). For Rényi, tightness on binary alphabets is verified for Df(PQ)=EQ[f(dPdQ)]D_f(P\|Q) = \mathbb{E}_Q\left[f\left(\frac{dP}{dQ}\right)\right]3 (Grosse et al., 20 Jan 2025).
  • Range of Validity: The permitted TV is bounded by a function of Df(PQ)=EQ[f(dPdQ)]D_f(P\|Q) = \mathbb{E}_Q\left[f\left(\frac{dP}{dQ}\right)\right]4, specifically, Df(PQ)=EQ[f(dPdQ)]D_f(P\|Q) = \mathbb{E}_Q\left[f\left(\frac{dP}{dQ}\right)\right]5, reflecting the feasible overlap of Radon–Nikodým ratios (Binette, 2018).

In the small TV regime, the minimal KL divergence one must pay to move at least Df(PQ)=EQ[f(dPdQ)]D_f(P\|Q) = \mathbb{E}_Q\left[f\left(\frac{dP}{dQ}\right)\right]6 in TV away from a reference measure Df(PQ)=EQ[f(dPdQ)]D_f(P\|Q) = \mathbb{E}_Q\left[f\left(\frac{dP}{dQ}\right)\right]7 is characterized by

Df(PQ)=EQ[f(dPdQ)]D_f(P\|Q) = \mathbb{E}_Q\left[f\left(\frac{dP}{dQ}\right)\right]8

with Df(PQ)=EQ[f(dPdQ)]D_f(P\|Q) = \mathbb{E}_Q\left[f\left(\frac{dP}{dQ}\right)\right]9 for "balanced" TV(P,Q)=12dP/dQ1dQ.\mathrm{TV}(P, Q) = \frac{1}{2} \int |dP/dQ - 1| \, dQ.0 and TV(P,Q)=12dP/dQ1dQ.\mathrm{TV}(P, Q) = \frac{1}{2} \int |dP/dQ - 1| \, dQ.1 for unbalanced TV(P,Q)=12dP/dQ1dQ.\mathrm{TV}(P, Q) = \frac{1}{2} \int |dP/dQ - 1| \, dQ.2 (with TV(P,Q)=12dP/dQ1dQ.\mathrm{TV}(P, Q) = \frac{1}{2} \int |dP/dQ - 1| \, dQ.3 the balance coefficient) (Berend et al., 2012).

5. Generalizations and Channel-Based Inequalities

  • Strong Data Processing Inequalities (SDPIs): Reverse-Pinsker-type upper bounds, when combined with Pinsker-type lower bounds, yield contraction inequalities for TV(P,Q)=12dP/dQ1dQ.\mathrm{TV}(P, Q) = \frac{1}{2} \int |dP/dQ - 1| \, dQ.4-divergences under the action of a fixed channel (Markov operator). Given a channel TV(P,Q)=12dP/dQ1dQ.\mathrm{TV}(P, Q) = \frac{1}{2} \int |dP/dQ - 1| \, dQ.5 and a set TV(P,Q)=12dP/dQ1dQ.\mathrm{TV}(P, Q) = \frac{1}{2} \int |dP/dQ - 1| \, dQ.6 of allowed priors,

TV(P,Q)=12dP/dQ1dQ.\mathrm{TV}(P, Q) = \frac{1}{2} \int |dP/dQ - 1| \, dQ.7

where TV(P,Q)=12dP/dQ1dQ.\mathrm{TV}(P, Q) = \frac{1}{2} \int |dP/dQ - 1| \, dQ.8 is the total variation contraction coefficient and TV(P,Q)=12dP/dQ1dQ.\mathrm{TV}(P, Q) = \frac{1}{2} \int |dP/dQ - 1| \, dQ.9 encodes the [Harremoës–Vajda] lower bound (Grosse et al., 20 Jan 2025).

  • Local Information Geometry: The generalized quasi-m=essinfQdP/dQm = \mathrm{ess\,inf}_Q \,dP/dQ0-neighborhood framework enables comparison of all m=essinfQdP/dQm = \mathrm{ess\,inf}_Q \,dP/dQ1-divergences locally via second and third derivatives of their generators (Yu et al., 2024). Within these neighborhoods, tight equivalences exist:

m=essinfQdP/dQm = \mathrm{ess\,inf}_Q \,dP/dQ2

and

m=essinfQdP/dQm = \mathrm{ess\,inf}_Q \,dP/dQ3

with analogous results for Hellinger and other metrics (Yu et al., 2024).

6. Applications, Implications, and Limitations

Reverse-Pinsker-type bounds play a central role in:

  • Statistical Inference: Establishing local equivalence between divergences and TV for rates of convergence, consistency, or hypothesis testing error bounds (Yu et al., 2024, Berend et al., 2012).
  • Large Deviation Theory: Sharpening rate functions, e.g., beyond the universal m=essinfQdP/dQm = \mathrm{ess\,inf}_Q \,dP/dQ4, with quadratic coefficients m=essinfQdP/dQm = \mathrm{ess\,inf}_Q \,dP/dQ5 reflecting the structure of m=essinfQdP/dQm = \mathrm{ess\,inf}_Q \,dP/dQ6 (Berend et al., 2012).
  • Privacy and Information Theory: In privacy amplification, post-processing of mechanisms leads to improved Rényi-local differential privacy guarantees (Grosse et al., 20 Jan 2025).

A key limitation is that, in general, TV alone cannot control divergence unless essential supremum and infimum of the likelihood ratio are bounded (otherwise, divergence can be arbitrarily large for fixed TV). All nontrivial reverse-Pinsker bounds require such support, boundedness, or neighborhood constraints (Binette, 2018, Sason, 2015).

7. Comparative Overview and Further Developments

Multiple lines of work have pursued increasingly tight or general forms:

  • Binette: Best-possible, support-constrained m=essinfQdP/dQm = \mathrm{ess\,inf}_Q \,dP/dQ7-divergence bounds (Binette, 2018).
  • Sason & Verdú: Sharp and general reverse-Pinsker for both KL and Rényi, including TV, extremal ratios, and refined bounds for finite alphabets (Sason, 2015).
  • Harremoës–Vajda: Joint range methods for both lower and upper bounds, emphasizing exact phase transitions in tightness (Grosse et al., 20 Jan 2025).
  • Berend–Harremoës–Kontorovich: For "minimum KL divergence at a given TV," identify explicit quadratic behavior and its dependence on the structure of m=essinfQdP/dQm = \mathrm{ess\,inf}_Q \,dP/dQ8 (Berend et al., 2012).
  • Functional Analytic and Quantum Generalizations: Operator-convexity and reduction to two-point (commutative) or projective reductions attain corresponding quantum bounds (Rastegin, 2011).

The interplay between global bounds (for large m=essinfQdP/dQm = \mathrm{ess\,inf}_Q \,dP/dQ9 or small M=esssupQdP/dQM = \mathrm{ess\,sup}_Q \,dP/dQ0), local geometric expansions, and optimality under support constraints represents a unifying theme across settings. These results clarify the possible relationships between divergence and TV, providing essential tools in diverse areas of information theory and statistics.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Reverse-Pinsker-Type Bound.