Reverse-Pinsker-Type Bound
- Reverse-Pinsker-Type Bound is a framework that provides sharp upper bounds on f-divergences using total variation distance and pointwise constraints on the likelihood ratio.
- The approach is applied to various divergences, including KL, Rényi, and χ², through convex functions, optimal three-point discrete models, and tight finite-support constructions.
- These bounds have practical significance in statistical inference, large deviation theory, and privacy, linking divergence metrics with TV for enhanced analytical and operational insights.
A reverse-Pinsker-type bound refers to a family of sharp upper bounds on -divergences—in particular, those generated by convex functions such as the Kullback–Leibler (KL), Rényi, or χ² divergences—in terms of the total variation (TV) distance between two probability measures, sometimes incorporating pointwise constraints on the Radon–Nikodým derivative. While the classical Pinsker inequality gives a lower bound on TV in terms of divergence ("direct Pinsker"), its reverse addresses the fundamental question of how large a divergence can be if only the TV distance and certain extremal values of likelihood ratio are specified. These inequalities are operationally significant in information theory, statistics, large deviations, and local information geometry.
1. Core Definitions and Framework
Let denote two probability measures on a measurable space, and let be a convex function with . The -divergence is
and the total variation distance is
For pointwise control, let and (alternatively, the essential extrema of the "relative information" ).
General Reverse-Pinsker-Type Bound (Binette (Binette, 2018)): For 0 and 1,
2
with equality achieved (optimality) for a discrete three-point model. This principle extends to many divergence types by choosing the appropriate convex generator 3.
2. Major Classes of Reverse-Pinsker Inequalities
Reverse-Pinsker-type bounds have been developed for a variety of divergences and settings:
- KL Divergence: For 4, the sharp reverse-Pinsker bound is
5
where 6, 7 are the essential bounds on the Radon–Nikodým derivative reciprocal (Binette, 2018).
- Rényi Divergence: For order 8, the optimal bound is
9
where 0, 1, 2, and
3
(Grosse et al., 20 Jan 2025). This is tight for TV distance 4.
- 5-divergence: For 6,
7
- Hellinger and General 8-divergences: Similar structure, with explicit formulas for Hellinger of order 9 and other divergences, all subsumed under the Binette framework (Binette, 2018, Grosse et al., 20 Jan 2025).
A summary of specializations is shown below:
| Divergence | Generator 0 | Explicit Bound |
|---|---|---|
| KL | 1 | See above |
| 2 | 3 | 4 |
| Rényi-5 | 6 | 7 |
| Hellinger | 8 | Explicit in 9 |
3. Methodological Foundations and Proof Techniques
The foundational method involves a variational maximization of the divergence under TV and support constraints, usually reducing the extremal case to finite support models. The proofs typically follow these steps:
- Decomposition: Partition the space into 0 and its complement.
- Convexity Application: Use convexity of 1 and properties such as Jensen’s inequality, often combined with tight supporting lines at 2.
- Explicit Model Construction: Build discrete measures (typically on two or three points) achieving the bound, confirming tightness.
- Operator methods: In the quantum or functional-analytic setting, operator-convexity replaces standard convexity (Rastegin, 2011).
- Alternative Proofs: For Rényi, 3 (hockey-stick) representations integrate chordal approximations of 4 weighted by the second derivative of 5 (Grosse et al., 20 Jan 2025).
For the local (small perturbation) regime, Taylor expansion of the generator 6 up to order three yields
7
for 8, with 9 a quadratic energy and 0 controlling the third derivative, so
1
4. Extensions, Optimality, and Range
Reverse-Pinsker-type inequalities are sharp:
- Optimality: For fixed TV and extremal ratios (2), the discrete three-point construction achieves equality (Binette, 2018). For Rényi, tightness on binary alphabets is verified for 3 (Grosse et al., 20 Jan 2025).
- Range of Validity: The permitted TV is bounded by a function of 4, specifically, 5, reflecting the feasible overlap of Radon–Nikodým ratios (Binette, 2018).
In the small TV regime, the minimal KL divergence one must pay to move at least 6 in TV away from a reference measure 7 is characterized by
8
with 9 for "balanced" 0 and 1 for unbalanced 2 (with 3 the balance coefficient) (Berend et al., 2012).
5. Generalizations and Channel-Based Inequalities
- Strong Data Processing Inequalities (SDPIs): Reverse-Pinsker-type upper bounds, when combined with Pinsker-type lower bounds, yield contraction inequalities for 4-divergences under the action of a fixed channel (Markov operator). Given a channel 5 and a set 6 of allowed priors,
7
where 8 is the total variation contraction coefficient and 9 encodes the [Harremoës–Vajda] lower bound (Grosse et al., 20 Jan 2025).
- Local Information Geometry: The generalized quasi-0-neighborhood framework enables comparison of all 1-divergences locally via second and third derivatives of their generators (Yu et al., 2024). Within these neighborhoods, tight equivalences exist:
2
and
3
with analogous results for Hellinger and other metrics (Yu et al., 2024).
6. Applications, Implications, and Limitations
Reverse-Pinsker-type bounds play a central role in:
- Statistical Inference: Establishing local equivalence between divergences and TV for rates of convergence, consistency, or hypothesis testing error bounds (Yu et al., 2024, Berend et al., 2012).
- Large Deviation Theory: Sharpening rate functions, e.g., beyond the universal 4, with quadratic coefficients 5 reflecting the structure of 6 (Berend et al., 2012).
- Privacy and Information Theory: In privacy amplification, post-processing of mechanisms leads to improved Rényi-local differential privacy guarantees (Grosse et al., 20 Jan 2025).
A key limitation is that, in general, TV alone cannot control divergence unless essential supremum and infimum of the likelihood ratio are bounded (otherwise, divergence can be arbitrarily large for fixed TV). All nontrivial reverse-Pinsker bounds require such support, boundedness, or neighborhood constraints (Binette, 2018, Sason, 2015).
7. Comparative Overview and Further Developments
Multiple lines of work have pursued increasingly tight or general forms:
- Binette: Best-possible, support-constrained 7-divergence bounds (Binette, 2018).
- Sason & Verdú: Sharp and general reverse-Pinsker for both KL and Rényi, including TV, extremal ratios, and refined bounds for finite alphabets (Sason, 2015).
- Harremoës–Vajda: Joint range methods for both lower and upper bounds, emphasizing exact phase transitions in tightness (Grosse et al., 20 Jan 2025).
- Berend–Harremoës–Kontorovich: For "minimum KL divergence at a given TV," identify explicit quadratic behavior and its dependence on the structure of 8 (Berend et al., 2012).
- Functional Analytic and Quantum Generalizations: Operator-convexity and reduction to two-point (commutative) or projective reductions attain corresponding quantum bounds (Rastegin, 2011).
The interplay between global bounds (for large 9 or small 0), local geometric expansions, and optimality under support constraints represents a unifying theme across settings. These results clarify the possible relationships between divergence and TV, providing essential tools in diverse areas of information theory and statistics.