Smoothed Max-Information in Theory and Applications
- Smoothed max-information is a worst-case measure that quantifies how far a joint distribution deviates from the product form by discounting rare events.
- It plays a critical role in adaptive data analysis and differential privacy by controlling generalization and post-selection bias through bounded output and additive composition.
- In quantum settings, it acts as a one-shot analogue of mutual information, converging to von Neumann mutual information in the i.i.d. limit and aiding channel simulation and privacy amplification.
Smoothed max-information is a worst-case information measure that quantifies how far a joint object is from a product form after allowing a small amount of smoothing. In adaptive data analysis, the relevant notion is the -approximate max-information , which upper-bounds how much conditioning on an algorithm’s output can inflate the probability of any dataset and thereby controls generalization and post-selection bias (Rogers et al., 2016). In one-shot quantum information theory, smooth max-information is defined through max-relative entropy and purified-distance smoothing, where it serves as a one-shot analogue of mutual information and converges to von Neumann mutual information in the i.i.d. limit (Ciganović et al., 2013).
1. Classical approximate max-information
For jointly distributed random variables and , with denoting independent draws from their marginals and , the exact max-information is the minimal such that
for every in the supports of and 0. Equivalently,
1
This is a worst-case information measure, more stringent than Shannon mutual information (Rogers et al., 2016).
Smoothed, or approximate, max-information discards a bad event of probability at most 2. It is defined by
3
A central consequence is the event inequality
4
whenever 5. This says that, up to a multiplicative factor 6 and additive slack 7, the joint law behaves like the product law.
For a randomized algorithm 8, one writes 9 and studies 0 either over all input distributions or only over product distributions 1. The notation 2 denotes the worst case over all distributions on 3, while 4 denotes the worst case over product distributions. A simple special case is bounded description length: if 5 always outputs a string of at most 6 bits, then for any 7,
8
so short output alone yields a smoothed max-information bound (Rogers et al., 2016).
2. Differential privacy, generalization, and post-selection inference
The 2016 study of adaptive hypothesis testing identifies bounded approximate max-information as the key condition enabling valid post-selection correction. For pure differential privacy, if 9 is 0-differentially private, then
1
and over product distributions
2
For approximate privacy, the main theorem states that if 3 is 4-differentially private with 5 and 6, then for
7
one has
8
but only for product-distributed inputs (Rogers et al., 2016).
This bound feeds directly into generalization. If 9, 0, and 1, then for any event 2,
3
Applied to a low-sensitivity output function 4 with sensitivity 5, the paper gives
6
This nearly recovers the optimal low-sensitivity-query generalization bound of Bassily et al. (2015).
The same mechanism yields explicit post-selection 7-value correction. If a data-dependent selection rule 8 satisfies 9, then
0
is a valid 1-value correction function. Operationally, the significance threshold must be shrunk by roughly a factor 2, with additive slack 3, to compensate for adaptivity (Rogers et al., 2016).
3. Composition properties and structural limitations
Approximate max-information composes additively. If 4 and 5 satisfy
6
for every fixed 7, then the adaptive composition 8 satisfies
9
This additive composition is one of the main reasons the measure functions as a “budget accounting” device for adaptive analyses (Rogers et al., 2016).
A central limitation is that the connection between 0-differential privacy and bounded smoothed max-information is genuinely distribution-dependent. The paper constructs algorithms 1 and 2 such that 3 has small max-information, each fixed 4 is 5-differentially private and has small max-information on independent inputs, yet the composition reconstructs the entire dataset with probability at least 6. Consequently,
7
for 8. This proves that 9-DP can come first in a composition with max-information-bounded algorithms, but not necessarily second if the composition is required to satisfy a nontrivial max-information bound (Rogers et al., 2016).
This lower bound also resolves a common misconception. Pure 0-DP implies max-information bounds uniformly over all input distributions, whereas approximate 1-DP implies bounded smoothed max-information only under the product-distribution assumption. The product hypothesis is therefore essential, not a technical artifact.
4. Quantum smooth max-information as a one-shot mutual information
In the finite-dimensional quantum setting, smooth max-information is defined from the max-relative entropy
2
together with smoothing in purified distance. For subnormalized states, the purified distance is
3
where
4
The 5-ball is 6 (Ciganović et al., 2013).
The 2013 quantum treatment distinguishes three unsmoothed 7-based analogues of mutual information: 8
9
0
These definitions coincide for ordinary relative entropy, but not for 1. Definitions 2 and 3 are bounded by 2, whereas definition 1 can be unbounded. The smoothed quantities are obtained by minimizing the corresponding unsmoothed definition over 3 (Ciganović et al., 2013).
The central structural result is that the smoothed variants are essentially equivalent up to additive logarithmic terms in the smoothing parameters. This yields approximate symmetry for definition 2, data processing for all smoothed variants under local CPTP maps, and chain rules relating smooth max-information to differences of smooth min- and max-entropies. In the i.i.d. regime, the quantum asymptotic equipartition property implies that for each of the three smoothed definitions,
4
Accordingly, smooth max-information is a one-shot generalization of von Neumann mutual information (Ciganović et al., 2013).
5. Operational meanings in channel simulation and privacy amplification
A channel-level version is obtained by evaluating state max-information on the Choi state of a channel and then smoothing in diamond norm. For a channel 5,
6
This quantity has an exact operational interpretation: for no-signalling assisted simulation,
7
where 8 is the least rounding constant making the right-hand side an integer logarithm. The corresponding asymptotic equipartition property gives
9
which implies the no-signalling assisted quantum reverse Shannon theorem (Fang et al., 2018).
In quantum privacy amplification, the relevant smoothing problem is the exponential behavior of
0
For i.i.d. states, the exact exponent of the minimal smoothing error is
1
Applied to 2 versus 3, this yields upper bounds on the asymptotic decay of insecurity in privacy amplification, and for rates 4 the upper and lower bounds match, giving the exact security exponent (Li et al., 2021).
These operational results show that smooth max-information is not merely a formal replacement for mutual information. It determines one-shot communication cost in channel simulation and the decay of residual correlation in finite-blocklength cryptographic tasks.
6. Inequalities, minimax methods, and computation
A minimax approach to one-shot entropy inequalities gives dimension-free, 5-explicit control of smoothed max-divergence. For 6,
7
and for 8, 9,
00
Because smoothed max-information is obtained by evaluating smoothed max-divergence against a product reference state, these inequalities immediately transfer to one-shot mutual-information-type quantities (Anshu et al., 2019).
Computation of the quantum smooth max-mutual information 01 was addressed in 2025 by an iterative SDP-based algorithm. For
02
the paper gives a two-step mountain-climbing procedure: an SDP for the current 03 in
04
followed by an SDP over 05 maximizing a slack variable 06 in
07
The method is accurate if, for all 08, the state 09 is positive definite; otherwise it provides an upper bound. The paper establishes primal and dual SDPs and proves strong duality (Popp et al., 9 Sep 2025).
7. Infinite-dimensional extensions and conceptual synthesis
For infinite-dimensional separable Hilbert spaces, smoothed conditional min- and max-entropies were extended using the same purified-distance ball
10
together with
11
These entropies retain smooth duality and admit an infinite-dimensional AEP under appropriate assumptions (Furrer et al., 2010).
The paper does not explicitly define infinite-dimensional smoothed max-information, but its framework suggests the standard one-shot definition
12
A natural implication is that finite-dimensional results can be transported by truncating to projected finite-dimensional subspaces and passing to the limit, since the paper proves that min- and max-entropies of projected states converge to the infinite-dimensional quantities (Furrer et al., 2010).
Across classical adaptive statistics, one-shot quantum information, channel simulation, and privacy amplification, the common role of smoothing is to regularize a worst-case max-divergence notion without abandoning one-shot control. In the classical setting the smoothing parameter discards rare bad events and yields eventwise control of dependence; in the quantum setting it minimizes over nearby states or channels in purified or diamond distance. This suggests that smoothed max-information is best understood as a robust max-divergence-to-product principle: it measures residual dependence in a form strong enough to support composition, finite-blocklength converses, operational simulation theorems, and adaptive-valid statistical inference.