Papers
Topics
Authors
Recent
Search
2000 character limit reached

Smoothed Max-Information in Theory and Applications

Updated 14 July 2026
  • Smoothed max-information is a worst-case measure that quantifies how far a joint distribution deviates from the product form by discounting rare events.
  • It plays a critical role in adaptive data analysis and differential privacy by controlling generalization and post-selection bias through bounded output and additive composition.
  • In quantum settings, it acts as a one-shot analogue of mutual information, converging to von Neumann mutual information in the i.i.d. limit and aiding channel simulation and privacy amplification.

Smoothed max-information is a worst-case information measure that quantifies how far a joint object is from a product form after allowing a small amount of smoothing. In adaptive data analysis, the relevant notion is the β\beta-approximate max-information Iβ(X;Z)I_\infty^\beta(X;Z), which upper-bounds how much conditioning on an algorithm’s output can inflate the probability of any dataset and thereby controls generalization and post-selection bias (Rogers et al., 2016). In one-shot quantum information theory, smooth max-information is defined through max-relative entropy and purified-distance smoothing, where it serves as a one-shot analogue of mutual information and converges to von Neumann mutual information in the i.i.d. limit (Ciganović et al., 2013).

1. Classical approximate max-information

For jointly distributed random variables XX and ZZ, with XZX \otimes Z denoting independent draws from their marginals and log=log2\log = \log_2, the exact max-information is the minimal kk such that

Pr[X=xZ=z]2kPr[X=x]\Pr[X=x\mid Z=z] \le 2^k \Pr[X=x]

for every x,zx,z in the supports of XX and Iβ(X;Z)I_\infty^\beta(X;Z)0. Equivalently,

Iβ(X;Z)I_\infty^\beta(X;Z)1

This is a worst-case information measure, more stringent than Shannon mutual information (Rogers et al., 2016).

Smoothed, or approximate, max-information discards a bad event of probability at most Iβ(X;Z)I_\infty^\beta(X;Z)2. It is defined by

Iβ(X;Z)I_\infty^\beta(X;Z)3

A central consequence is the event inequality

Iβ(X;Z)I_\infty^\beta(X;Z)4

whenever Iβ(X;Z)I_\infty^\beta(X;Z)5. This says that, up to a multiplicative factor Iβ(X;Z)I_\infty^\beta(X;Z)6 and additive slack Iβ(X;Z)I_\infty^\beta(X;Z)7, the joint law behaves like the product law.

For a randomized algorithm Iβ(X;Z)I_\infty^\beta(X;Z)8, one writes Iβ(X;Z)I_\infty^\beta(X;Z)9 and studies XX0 either over all input distributions or only over product distributions XX1. The notation XX2 denotes the worst case over all distributions on XX3, while XX4 denotes the worst case over product distributions. A simple special case is bounded description length: if XX5 always outputs a string of at most XX6 bits, then for any XX7,

XX8

so short output alone yields a smoothed max-information bound (Rogers et al., 2016).

2. Differential privacy, generalization, and post-selection inference

The 2016 study of adaptive hypothesis testing identifies bounded approximate max-information as the key condition enabling valid post-selection correction. For pure differential privacy, if XX9 is ZZ0-differentially private, then

ZZ1

and over product distributions

ZZ2

For approximate privacy, the main theorem states that if ZZ3 is ZZ4-differentially private with ZZ5 and ZZ6, then for

ZZ7

one has

ZZ8

but only for product-distributed inputs (Rogers et al., 2016).

This bound feeds directly into generalization. If ZZ9, XZX \otimes Z0, and XZX \otimes Z1, then for any event XZX \otimes Z2,

XZX \otimes Z3

Applied to a low-sensitivity output function XZX \otimes Z4 with sensitivity XZX \otimes Z5, the paper gives

XZX \otimes Z6

This nearly recovers the optimal low-sensitivity-query generalization bound of Bassily et al. (2015).

The same mechanism yields explicit post-selection XZX \otimes Z7-value correction. If a data-dependent selection rule XZX \otimes Z8 satisfies XZX \otimes Z9, then

log=log2\log = \log_20

is a valid log=log2\log = \log_21-value correction function. Operationally, the significance threshold must be shrunk by roughly a factor log=log2\log = \log_22, with additive slack log=log2\log = \log_23, to compensate for adaptivity (Rogers et al., 2016).

3. Composition properties and structural limitations

Approximate max-information composes additively. If log=log2\log = \log_24 and log=log2\log = \log_25 satisfy

log=log2\log = \log_26

for every fixed log=log2\log = \log_27, then the adaptive composition log=log2\log = \log_28 satisfies

log=log2\log = \log_29

This additive composition is one of the main reasons the measure functions as a “budget accounting” device for adaptive analyses (Rogers et al., 2016).

A central limitation is that the connection between kk0-differential privacy and bounded smoothed max-information is genuinely distribution-dependent. The paper constructs algorithms kk1 and kk2 such that kk3 has small max-information, each fixed kk4 is kk5-differentially private and has small max-information on independent inputs, yet the composition reconstructs the entire dataset with probability at least kk6. Consequently,

kk7

for kk8. This proves that kk9-DP can come first in a composition with max-information-bounded algorithms, but not necessarily second if the composition is required to satisfy a nontrivial max-information bound (Rogers et al., 2016).

This lower bound also resolves a common misconception. Pure Pr[X=xZ=z]2kPr[X=x]\Pr[X=x\mid Z=z] \le 2^k \Pr[X=x]0-DP implies max-information bounds uniformly over all input distributions, whereas approximate Pr[X=xZ=z]2kPr[X=x]\Pr[X=x\mid Z=z] \le 2^k \Pr[X=x]1-DP implies bounded smoothed max-information only under the product-distribution assumption. The product hypothesis is therefore essential, not a technical artifact.

4. Quantum smooth max-information as a one-shot mutual information

In the finite-dimensional quantum setting, smooth max-information is defined from the max-relative entropy

Pr[X=xZ=z]2kPr[X=x]\Pr[X=x\mid Z=z] \le 2^k \Pr[X=x]2

together with smoothing in purified distance. For subnormalized states, the purified distance is

Pr[X=xZ=z]2kPr[X=x]\Pr[X=x\mid Z=z] \le 2^k \Pr[X=x]3

where

Pr[X=xZ=z]2kPr[X=x]\Pr[X=x\mid Z=z] \le 2^k \Pr[X=x]4

The Pr[X=xZ=z]2kPr[X=x]\Pr[X=x\mid Z=z] \le 2^k \Pr[X=x]5-ball is Pr[X=xZ=z]2kPr[X=x]\Pr[X=x\mid Z=z] \le 2^k \Pr[X=x]6 (Ciganović et al., 2013).

The 2013 quantum treatment distinguishes three unsmoothed Pr[X=xZ=z]2kPr[X=x]\Pr[X=x\mid Z=z] \le 2^k \Pr[X=x]7-based analogues of mutual information: Pr[X=xZ=z]2kPr[X=x]\Pr[X=x\mid Z=z] \le 2^k \Pr[X=x]8

Pr[X=xZ=z]2kPr[X=x]\Pr[X=x\mid Z=z] \le 2^k \Pr[X=x]9

x,zx,z0

These definitions coincide for ordinary relative entropy, but not for x,zx,z1. Definitions 2 and 3 are bounded by x,zx,z2, whereas definition 1 can be unbounded. The smoothed quantities are obtained by minimizing the corresponding unsmoothed definition over x,zx,z3 (Ciganović et al., 2013).

The central structural result is that the smoothed variants are essentially equivalent up to additive logarithmic terms in the smoothing parameters. This yields approximate symmetry for definition 2, data processing for all smoothed variants under local CPTP maps, and chain rules relating smooth max-information to differences of smooth min- and max-entropies. In the i.i.d. regime, the quantum asymptotic equipartition property implies that for each of the three smoothed definitions,

x,zx,z4

Accordingly, smooth max-information is a one-shot generalization of von Neumann mutual information (Ciganović et al., 2013).

5. Operational meanings in channel simulation and privacy amplification

A channel-level version is obtained by evaluating state max-information on the Choi state of a channel and then smoothing in diamond norm. For a channel x,zx,z5,

x,zx,z6

This quantity has an exact operational interpretation: for no-signalling assisted simulation,

x,zx,z7

where x,zx,z8 is the least rounding constant making the right-hand side an integer logarithm. The corresponding asymptotic equipartition property gives

x,zx,z9

which implies the no-signalling assisted quantum reverse Shannon theorem (Fang et al., 2018).

In quantum privacy amplification, the relevant smoothing problem is the exponential behavior of

XX0

For i.i.d. states, the exact exponent of the minimal smoothing error is

XX1

Applied to XX2 versus XX3, this yields upper bounds on the asymptotic decay of insecurity in privacy amplification, and for rates XX4 the upper and lower bounds match, giving the exact security exponent (Li et al., 2021).

These operational results show that smooth max-information is not merely a formal replacement for mutual information. It determines one-shot communication cost in channel simulation and the decay of residual correlation in finite-blocklength cryptographic tasks.

6. Inequalities, minimax methods, and computation

A minimax approach to one-shot entropy inequalities gives dimension-free, XX5-explicit control of smoothed max-divergence. For XX6,

XX7

and for XX8, XX9,

Iβ(X;Z)I_\infty^\beta(X;Z)00

Because smoothed max-information is obtained by evaluating smoothed max-divergence against a product reference state, these inequalities immediately transfer to one-shot mutual-information-type quantities (Anshu et al., 2019).

Computation of the quantum smooth max-mutual information Iβ(X;Z)I_\infty^\beta(X;Z)01 was addressed in 2025 by an iterative SDP-based algorithm. For

Iβ(X;Z)I_\infty^\beta(X;Z)02

the paper gives a two-step mountain-climbing procedure: an SDP for the current Iβ(X;Z)I_\infty^\beta(X;Z)03 in

Iβ(X;Z)I_\infty^\beta(X;Z)04

followed by an SDP over Iβ(X;Z)I_\infty^\beta(X;Z)05 maximizing a slack variable Iβ(X;Z)I_\infty^\beta(X;Z)06 in

Iβ(X;Z)I_\infty^\beta(X;Z)07

The method is accurate if, for all Iβ(X;Z)I_\infty^\beta(X;Z)08, the state Iβ(X;Z)I_\infty^\beta(X;Z)09 is positive definite; otherwise it provides an upper bound. The paper establishes primal and dual SDPs and proves strong duality (Popp et al., 9 Sep 2025).

7. Infinite-dimensional extensions and conceptual synthesis

For infinite-dimensional separable Hilbert spaces, smoothed conditional min- and max-entropies were extended using the same purified-distance ball

Iβ(X;Z)I_\infty^\beta(X;Z)10

together with

Iβ(X;Z)I_\infty^\beta(X;Z)11

These entropies retain smooth duality and admit an infinite-dimensional AEP under appropriate assumptions (Furrer et al., 2010).

The paper does not explicitly define infinite-dimensional smoothed max-information, but its framework suggests the standard one-shot definition

Iβ(X;Z)I_\infty^\beta(X;Z)12

A natural implication is that finite-dimensional results can be transported by truncating to projected finite-dimensional subspaces and passing to the limit, since the paper proves that min- and max-entropies of projected states converge to the infinite-dimensional quantities (Furrer et al., 2010).

Across classical adaptive statistics, one-shot quantum information, channel simulation, and privacy amplification, the common role of smoothing is to regularize a worst-case max-divergence notion without abandoning one-shot control. In the classical setting the smoothing parameter discards rare bad events and yields eventwise control of dependence; in the quantum setting it minimizes over nearby states or channels in purified or diamond distance. This suggests that smoothed max-information is best understood as a robust max-divergence-to-product principle: it measures residual dependence in a form strong enough to support composition, finite-blocklength converses, operational simulation theorems, and adaptive-valid statistical inference.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Smoothed Max-Information.