---
title: Quantum Smooth Max-Mutual Information
url: https://www.emergentmind.com/topics/quantum-smooth-max-mutual-information
type: topic
---

# Quantum Smooth Max-Mutual Information

Quantum smooth max-mutual information is a family of one-shot, non-asymptotic correlation measures obtained by replacing the relative entropy in quantum mutual information with the max-relative entropy \(D_{\max}\), and then smoothing over nearby states or channels. In contrast to the von Neumann case, the state-level max-information is not uniquely defined: several inequivalent unsmoothed formulas coexist, and a substantial part of the literature is devoted to showing that their smoothed versions are equivalent up to additive logarithmic terms in the smoothing parameters. The quantity has become central in one-shot quantum Shannon theory, where it appears in converse bounds for quantum state redistribution, in exact simulation-cost formulas for channels, and in cryptographic leakage bounds; in asymptotic iid limits, it reduces to ordinary mutual-information-type rates [1308.5884] [1409.4338] [1807.05354].

## 1. Definitions and variants

The common starting point is the max-relative entropy
\[
D_{\max}(\rho\|\sigma) := \log \min\{\lambda:\rho\le \lambda \sigma\},
\]
introduced as a one-shot relative-entropy quantity and later generalized to smooth versions [0803.2770].

At the bipartite state level, one of the earliest mutual-information analogues is
\[
D_{\max}(A:B) := D_{\max}(\rho_{AB}\,\|\,\rho_A\otimes \rho_B),
\]
with smoothed version
\[
D_{\max}^{\varepsilon}(A:B) := D_{\max}^{\varepsilon}(\rho_{AB}\,\|\,\rho_A\otimes \rho_B).
\]
This construction is explicitly presented as a one-shot analogue of mutual information in the Smooth Rényi Entropy framework, and its asymptotic behavior is linked to spectral mutual informations through the information-spectrum formalism [0803.2770].

A later systematic treatment distinguishes three state-level max-information functionals:
\[
{}^{1}I_{\max}(A:B)_\rho := D_{\max}(\rho_{AB}\|\rho_A\otimes \rho_B),
\]
\[
{}^{2}I_{\max}(A:B)_\rho := \min_{\sigma_B} D_{\max}(\rho_{AB}\|\rho_A\otimes \sigma_B),
\]
\[
{}^{3}I_{\max}(A:B)_\rho := \min_{\sigma_A,\sigma_B} D_{\max}(\rho_{AB}\|\sigma_A\otimes \sigma_B).
\]
The paper introducing this tripartite taxonomy emphasizes that these are all natural \(D_{\max}\)-based analogues of mutual information, but unlike the von Neumann case they are not exactly equivalent. In particular, \({}^{1}I_{\max}\) is generally unbounded, whereas \({}^{2}I_{\max}\) and \({}^{3}I_{\max}\) satisfy
\[
{}^{2}I_{\max},\,{}^{3}I_{\max}\ \le\ 2\log\min\{|A|,|B|\}
\]
on finite-dimensional systems [1308.5884].

Operational papers often adopt the optimized-over-\(B\) convention
\[
I_{\max}(A:B)_\rho := \inf_{\sigma^B\in D_=(B)} D_{\max}\!\big(\rho^{AB}\,\big\|\,\rho^A\otimes \sigma^B\big),
\]
which is the form used in one-shot state redistribution converses [1409.4338]. A cryptographic line of work uses a closely related smoothed quantity
\[
I_{\max}^{\epsilon}(A;B)_\rho
=
\inf_{\substack{\tilde\rho\in S_{\le}(AB)\\ P(\tilde\rho,\rho)\le \epsilon}}
\inf_{\sigma_B\in S_{\le}(B)}
D_{\max}(\tilde\rho_{AB}\| \tilde\rho_A\otimes \sigma_B),
\]
and explicitly notes that, for normalized \(\rho\), the smoothing can be taken over normalized \(\tilde\rho\) for this variant [2407.20396].

A more recent computational convention fixes the original first marginal and lets only the second marginal change with the smoothed state:
\[
I_{\max}^\varepsilon(\rho_{AB}) :=
\inf_{\tilde{\rho}_{AB}\in B^\varepsilon(\rho_{AB})}
D_{\max}\!\bigl(\tilde{\rho}_{AB}\,\|\,\rho_A\otimes \tilde{\rho}_B\bigr),
\qquad
\tilde{\rho}_B:=\Tr_A[\tilde{\rho}_{AB}],
\]
which is the definition used in the semidefinite-programming approach of 2025 [2509.07743].

The same replacement principle extends to channels. For a quantum channel \(\mathcal N_{A'\to B}\), the channel’s max-information is defined by
\[
I_{\max}(A:B)_{\mathcal N} := I_{\max}(A:B)_{\mathcal N_{A'\to B}(\Phi_{AA'})},
\]
and its smoothed version is obtained by optimizing over channels within a diamond-norm ball around \(\mathcal N\) [1807.05354].

## 2. Smoothing conventions, geometry, and equivalence

The literature uses several distinct smoothing geometries. Early work on \(D_{\max}^{\varepsilon}\) smooths over a trace-norm ball of subnormalized states,
\[
B_\varepsilon(\rho) :=
\left\{\bar\rho\ge 0:\ \|\bar\rho-\rho\|_1\le \varepsilon,\ \mathrm{Tr}\,\bar\rho \le \mathrm{Tr}\,\rho\right\},
\]
which is explicitly non-asymptotic and one-shot [0803.2770]. A later mutual-information-focused treatment adopts purified distance
\[
P(\rho,\sigma):=\sqrt{1-F(\rho,\sigma)^2},
\]
with \(\varepsilon\)-balls of subnormalized states, and uses this as the common smoothing metric for all three max-information definitions [1308.5884].

In one-shot state redistribution, smoothing is also performed with purified distance, but the paper stresses a point that is technically important for smooth max-information bounds: the optimization is over all nearby subnormalized states in purified distance,
\[
B^\varepsilon(\rho):=\{\bar\rho\in D_\le(A): P(\rho,\bar\rho)\le \varepsilon\},
\]
rather than by perturbing marginals separately. The same discussion explicitly emphasizes that the relevant smooth entropies are smoothed simultaneously on overlapping systems, even though that is generally subtle [1409.4338].

Other conventions coexist. A 2025 Rényi-based paper defines
\[
I_{\max}^{\varepsilon}(A:B)_\rho
=
\min_{\rho' \in B^\varepsilon(\rho^{AB})} I_{\max}(A:B)_{\rho'}
\]
with \(B^\varepsilon\) a trace-distance ball [2502.06526]. The 2025 SDP paper uses the fidelity-induced sine distance
\[
P(\rho,\sigma):=\sqrt{1-\mathcal{F}(\rho,\sigma)},
\qquad
B^\varepsilon(\rho)=\{\tilde\rho:\mathcal{F}(\rho,\tilde\rho)\ge 1-\varepsilon^2\},
\]
and restricts the smoothing domain to density operators [2509.07743]. At the channel level, smoothing is over channels \(\widetilde{\mathcal N}\) satisfying
\[
\frac12\|\widetilde{\mathcal N}-\mathcal N\|_\diamond \le \varepsilon
\]
[1807.05354].

A central structural result is that the three smoothed state-level definitions are essentially equivalent up to additive correction terms depending only on the smoothing parameters. Specifically,
\[
{}^{3}I_{\max}^{\varepsilon+\varepsilon'}(A:B)_\rho
\le
{}^{2}I_{\max}^{\varepsilon}(A:B)_\rho
\le
{}^{3}I_{\max}^{\varepsilon}(A:B)_\rho + f(\varepsilon,\varepsilon'),
\]
and
\[
{}^{2}I_{\max}^{\varepsilon+2\sqrt{\varepsilon}+\varepsilon'}(A:B)_\rho
\le
{}^{1}I_{\max}^{\varepsilon+2\sqrt{\varepsilon}+\varepsilon'}(A:B)_\rho
\le
{}^{2}I_{\max}^{\varepsilon}(A:B)_\rho + g(\varepsilon),
\]
with \(f\) and \(g\) depending only on the smoothing parameters. This yields pairwise equivalence of all three smoothed notions and implies approximate symmetry of \({}^{2}I_{\max}\) up to smoothing corrections [1308.5884].

The joint-smoothing problem is addressed from a different angle by a minimax approach to one-shot entropy inequalities. That work does not define smooth max-mutual information directly, but it develops the divergence tools typically used to construct it. In particular, it proves a simultaneous smoothing theorem: for \(\rho_{AB}\), arbitrary \(\sigma_A,\sigma_B\), and \(\varepsilon+\varepsilon'<1\), there exists \(\tilde\rho_{AB}\) with
\[
P(\rho_{AB},\tilde\rho_{AB})\le \sqrt{\varepsilon+\varepsilon'}
\]
such that
\[
D_{\max}(\tilde\rho_A\|\sigma_A) \le D_h^{1-\varepsilon}(\rho_A\|\sigma_A)+\Delta,
\qquad
D_{\max}(\tilde\rho_B\|\sigma_B) \le D_h^{1-\varepsilon'}(\rho_B\|\sigma_B)+\Delta,
\]
where \(\Delta=-\log(1-\varepsilon-\varepsilon')\). This provides a single nearby bipartite state whose two marginals are simultaneously controlled, a structure that is directly relevant to smooth max-information manipulations [1906.00333].

## 3. One-shot state redistribution and converse bounds

In quantum state redistribution there are four systems of interest: \(A\) held by Alice, \(B\) held by Bob, \(C\) to be transmitted from Alice to Bob, and \(R\) purifying the \(ABC\) state. The main one-shot achievability theorem in this setting is expressed in terms of smooth conditional min- and max-entropies rather than directly in terms of \(I_{\max}\), but smooth max-information enters explicitly on the converse side [1409.4338].

For every one-shot state redistribution protocol \(\Pi\) with error \(\varepsilon_1\), and any \(\varepsilon_2\in(0,1-\varepsilon_1)\), the communication cost satisfies
\[
q(\Pi)\ge \frac12 I_{\max}^{\varepsilon_1+\varepsilon_2}(R;BC)_\rho -\frac12 I_{\max}^{\varepsilon_2}(R;B)_\rho.
\]
Equivalent converse lower bounds are given in terms of smooth conditional min- and max-entropies, and the same inequalities hold with \(B\) replaced by \(A\). This is the paper’s explicit operational appearance of smooth max-information: the unavoidable quantum communication is lower-bounded by a difference of two smooth max-information terms [1409.4338].

The same lower bound survives arbitrary back-communication. For feedback protocols,
\[
QCC_{A\to B}(\Pi)\ge \frac12 I_{\max}^{\varepsilon_1+\varepsilon_2}(R;BC)_\rho -\frac12 I_{\max}^{\varepsilon_2}(R;B)_\rho,
\]
so the amount of communication from Alice to Bob remains controlled by the same smooth max-information difference. This is the mechanism behind the strong converse statement in the interactive setting [1409.4338].

A key technical ingredient is the dimension blow-up lemma
\[
I_{\max}^{\varepsilon}(A:BC)_\rho\le I_{\max}^{\varepsilon}(A:B)_\rho+2\log|C|,
\]
which is used in the converse proof to peel off communicated registers one by one. Each transmitted qubit register contributes at most \(2\log|C_i|\) to the max-information [1409.4338].

The same paper introduces, in its concluding remarks, a conditional max-information candidate,
\[
I_{\max}(C;R|B)_\rho
:=
D_{\max}\!\Big(\rho^{CBR}\,\Big\|\,(\rho^{BR})^{1/2}(\rho^{B})^{-1/2}\rho^{BC}(\rho^{B})^{-1/2}(\rho^{BR})^{1/2}\Big),
\]
as a possible route to sharpening one-shot achievability bounds via Rényi-type quantities, but it is not used in the main theorem statements. The suggested improved rate estimate,
\[
\frac12 \max_i\!\left[H_{\max}(C|B)_{\rho_i}-H_{\min}(C|BR)_{\rho_i}\right],
\]
is left open [1409.4338].

In the iid asymptotic limit, the one-shot upper and lower bounds converge to the standard redistribution rates:
\[
q \to \frac12 I(C;R|B),
\qquad
e+q \to H(C|B).
\]
More precisely, the achievability side yields protocols with exponentially small error and
\[
q(\Pi_n)\le n\left[\frac12 I(C;R|B)+\mu\right],
\]
while the converse gives, for error \(1-2^{-cn}\),
\[
q(\Pi_n)\ge n\left[\frac12 I(C;R|B)-\mu\right].
\]
Thus the smooth max-information converse converges to the conditional mutual information rate and establishes a strong converse [1409.4338].

## 4. Chain rules, Rényi bounds, and convex-split refinements

One of the main motivations for smooth max-information is to extend the smooth entropy formalism from entropies to mutual-information-like quantities. In that direction, a foundational result is that smoothed max-information satisfies chain rules involving smooth min- and max-entropies. For example,
\[
{}^{2}I_{\max}^{\varepsilon}(A:B)_\rho \ge H_{\min}^{\varepsilon}(A)_\rho - H_{\min}^{\varepsilon}(A|B)_\rho,
\]
and the same framework provides upper chain rules with logarithmic smoothing corrections, thereby making max-information usable as a bridge between one-shot correlation measures and entropy-based techniques [1308.5884].

A different line of development derives direct Rényi-type bounds on smoothed max-information. A 2025 paper defines
\[
I_\alpha(A:B)_\rho \coloneqq \min_{\sigma\in D(B)} D_\alpha\!\left(\rho^{AB}\big\|\rho^A\otimes \sigma^B\right),
\]
so that \(I_{\max}\) is the special case \(\alpha=\infty\), and proves the dimension-independent universal upper bound
\[
I_{\max}^{\varepsilon}(A:B)_\rho \le H_\alpha(A)_\rho - H_\beta^{\varepsilon}(A|B)_\rho + f_{\alpha,\beta}(\varepsilon),
\]
valid for every \(\rho^{AB}\), every \(\alpha\in(0,1)\), every \(\beta>1\), and every \(\varepsilon\in(0,1)\). The correction term is
\[
f_{\alpha,\beta}(\varepsilon)
=
g(\alpha,\beta) + \left(\frac{4}{\beta-1}+\frac{2}{1-\alpha}\right)\frac{1}{\varepsilon},
\]
with
\[
g(\alpha,\beta)
=
\frac{(1/c)}{1-\alpha}
+\frac{(2/c^2)}{\beta-1}
+\frac{1}{2(1-c^2)},
\qquad
c = 2-\sqrt{3}\approx 0.27.
\]
The paper calls this bound “universal” because it depends only on Rényi entropies and the smoothing parameters, not on system dimension [2502.06526].

The same work refines the convex split lemma by replacing max-mutual information with collision mutual information. Instead of the standard inequality involving \(D_{\max}\), it proves an exact identity for the convex-split state,
\[
Q_2\!\left(\tau^{RA^n}\big\|\omega^R\otimes(\sigma^A)^{\otimes n}\right)
=
\frac{n-1}{n}Q_2(\rho^R\|\omega^R)
+
\frac1n Q_2(\rho^{RA}\|\omega^R\otimes\sigma^A).
\]
With \(\omega^R=\rho^R\), this becomes
\[
D_2\!\left(\tau^{RA^n}\big\|\rho^R\otimes(\sigma^A)^{\otimes n}\right) = 1+\frac{\mu}{n}.
\]
This replacement of a max-information inequality by a collision-information equality is stated to yield tighter achievability bounds for state splitting, state merging, and state redistribution, and to sharpen finite-blocklength bounds in reverse quantum Shannon simulation [2502.06526].

A complementary technical route is provided by the minimax framework for one-shot entropy inequalities. That work proves
\[
D_{\max}^{\varepsilon,P}(\rho\|\sigma)
\le
\widetilde D_\alpha(\rho\|\sigma)
+
\frac{1}{\alpha-1}\log\frac{1}{\varepsilon^2}
+
\log\frac{1}{1-\varepsilon^2},
\]
for \(\alpha>1\), and presents this as one of the key ingredients used to bound smooth max-mutual information through optimized product-reference divergences. The same paper also derives two-sided comparisons between smooth max-divergence and hypothesis testing divergence, as well as a joint-smoothing theorem for bipartite states [1906.00333].

## 5. Channel smooth max-information and simulation cost

For channels, smooth max-information acquires an exact one-shot operational meaning. Given a channel \(\mathcal N_{A'\to B}\), its max-information is defined by evaluating the state max-information on \(\mathcal N(\Phi_{AA'})\), and the smoothed version optimizes over channels \(\widetilde{\mathcal N}\) within diamond distance \(\varepsilon\) of \(\mathcal N\):
\[
I_{\max}^{\varepsilon}(A:B)_{\mathcal N}
:=
\inf_{\substack{\frac12\|\widetilde{\mathcal N}-\mathcal N\|_\diamond \le \varepsilon\\ \widetilde{\mathcal N}\in \mathrm{CPTP}(A':B)}}
I_{\max}(A:B)_{\widetilde{\mathcal N}}.
\]
The paper positions this as the one-shot generalization of channel mutual information [1807.05354].

Its main theorem identifies the one-shot \(\varepsilon\)-error NS-assisted quantum simulation cost exactly:
\[
S^{(1)}_{\mathrm{NS},\varepsilon}(\mathcal N)
=
\frac12 I_{\max}^{\varepsilon}(A:B)_{\mathcal N} + \delta,
\]
where \(\delta\in[0,1]\) is the least correction making the right-hand side the logarithm of an integer. This gives the channel’s smooth max-information an exact operational interpretation as the one-shot quantum simulation cost under no-signalling assisted codes [1807.05354].

The same framework admits equivalent resource-theoretic descriptions. If \(\boldsymbol{\mathcal G}\) denotes the set of constant channels, then
\[
I_{\max}^{\varepsilon}(A:B)_{\mathcal N}
=
\min_{\mathcal M\in \boldsymbol{\mathcal G}} D_{\max}^{\varepsilon}(\mathcal N\|\mathcal M),
\]
so the quantity may be viewed as a smoothed distance from useless channels. In terms of smoothed generalized robustness,
\[
I_{\max}^{\varepsilon}(A:B)_{\mathcal N}
=
\log\bigl(1+\mathcal R_g^\varepsilon(\mathcal N)\bigr).
\]
The paper also remarks that data processing under superchannels should hold, reflecting the intuition that noisier channels should not require more simulation resources [1807.05354].

An asymptotic equipartition property is established at the channel level:
\[
\lim_{\varepsilon\to 0}\lim_{n\to\infty}\frac1n I_{\max}^{\varepsilon}(A:B)_{\mathcal N^{\otimes n}}
=
I(A:B)_{\mathcal N}.
\]
This directly yields the no-signalling-assisted quantum reverse Shannon theorem. The paper also provides closed-form zero-error NS-assisted simulation costs for several canonical channels, including depolarizing, amplitude-damping, dephasing, and erasure channels, and notes that depolarizing and erasure channels have the same zero-error NS-assisted simulation cost [1807.05354].

## 6. Computation, cryptographic leakage, and current directions

Although smooth max-mutual information is defined by nested optimizations, it can now be addressed computationally. A 2025 paper presents an iterative SDP-based algorithm for
\[
I_{\max}^\varepsilon(\rho_{AB})
=
\inf_{\tilde{\rho}_{AB}\in B^\varepsilon(\rho_{AB})}
D_{\max}\!\bigl(\tilde{\rho}_{AB}\,\|\,\rho_A\otimes \tilde{\rho}_B\bigr),
\]
where smoothing is over a fidelity ball. The authors isolate the bilinear term \(\lambda(\rho_A\otimes\tilde\rho_B)\) as the main obstacle to a single SDP formulation, introduce an auxiliary SDP for the smoothing-update step, derive primal and dual forms, and prove strong duality. The overall method is a seesaw or mountain-climbing iteration: it is exact if, for \(\rho_{AB}\) and all \(\tilde\rho_{AB}\in B^\varepsilon(\rho_{AB})\), the product \(\rho_A\otimes\tilde\rho_B\) is positive definite; otherwise it returns an upper bound rather than necessarily the exact optimum [2509.07743].

In quantum cryptography, smooth max-information functions as a leakage measure. For a state \(\rho_{SLE}\), where \(S\) is the secret, \(E\) is side information, and \(L\) is an additional leakage register, a central chain rule states
\[
H_{\min}^{\epsilon+\delta}(S|LE)_\rho
\ge
H_{\min}^{\epsilon}(S|E)_\rho
-
I_{\max}^{\delta}(SE;L)_\rho
-
\log\!\left(\frac{4}{\delta^2}\right).
\]
A direct non-smoothed variant is
\[
H_{\min}(S|LE)_\rho \ge H_{\min}(S|E)_\rho - I_{\max}(L:SE)_\rho.
\]
The significance emphasized in this line of work is that the leakage penalty is expressed through a correlation quantity, not merely through the dimension of \(L\) [2407.20396].

The same paper derives an “information bounding theorem” for multi-round leakage processes. For a sequence of channels \(\{\mathcal M_i:R_{i-1}\to R_iL_i\}_{i=1}^n\),
\[
I_{\max}^{\epsilon}(SE:L_1^n)_{\mathcal M_n\circ\cdots\circ \mathcal M_1(\rho)}
\le
\sum_{i=1}^n
\sup_{\sigma\in S_{=}(R_{i-1}SE)}
I^\uparrow_\alpha(L_1^{i-1}SE;L_i)_{\mathcal M_i(\sigma)}
-
\frac{g(\epsilon)}{\alpha-1}.
\]
This places smooth max-information into an entropy-accumulation-style framework for handling leakage registers in QKD, randomness generation, and other security proofs robust against device imperfections [2407.20396].

A broader historical point is that the asymptotic interpretation of smooth max-type quantities was already present in the original max-relative-entropy framework: the smooth max-relative entropy converges to the sup-spectral divergence rate,
\[
\overline{D}(\rho\|\sigma)
=
\lim_{\varepsilon\to 0}\ \limsup_{n\to\infty}\ \frac{1}{n} D_{\max}^{\varepsilon}(\rho_n\|\sigma_n),
\]
and spectral mutual informations are recovered by substituting \(\sigma_n=\rho_{A,n}\otimes\rho_{B,n}\). This establishes the one-shot-to-asymptotic bridge on which later mutual-information and channel-information results build [0803.2770].

Taken together, these developments show that quantum smooth max-mutual information is not a single formula but a tightly connected cluster of one-shot correlation measures. The main open texture of the subject lies not in whether such a quantity exists, but in how different smoothing conventions, optimization domains, and operational tasks select different representatives of the same broader max-information paradigm.

Source: https://www.emergentmind.com/topics/quantum-smooth-max-mutual-information