---
title: Smoothed Quantum Divergences
url: https://www.emergentmind.com/topics/smoothed-quantum-divergences
type: topic
---

# Smoothed Quantum Divergences

Searching arXiv for recent papers on smoothed quantum divergences and closely related frameworks.
Smoothed quantum divergences are quantum divergences equipped with an error tolerance parameter that optimizes the divergence over states within a prescribed neighborhood, typically a trace-distance or purified-distance ball. In one-shot quantum information theory, this smoothing converts exact distinguishability measures into robust finite-blocklength quantities that control operational tasks such as hypothesis testing, channel coding, privacy amplification, state splitting, state merging, quantum compression, and resource conversion. Recent work has expanded the notion beyond standard trace-distance smoothing to include induced divergences derived from a parent relative entropy, universal order-to-order comparison bounds, and partial smoothing with fixed marginal constraints, thereby clarifying both the geometry and the operational role of smoothing in quantum information [2502.13669], [2601.10395], [2603.09885], [1807.05630].

## 1. Formal definitions and principal variants

A standard trace-distance formulation begins with the trace distance
\[
T(\rho,\sigma)\;:=\;\tfrac12\|\rho-\sigma\|_1
\]
and the \(\varepsilon\)-smoothing ball
\[
B^\varepsilon(\rho)\;:=\;\{\rho'\in S(\mathcal H)\;|\;T(\rho,\rho')\le\varepsilon\}.
\]
For any divergence \(D(\rho\|\sigma)\), the \(\varepsilon\)-smoothed divergence is defined by
\[
D^\varepsilon(\rho\|\sigma)\;:=\;\inf_{\rho'\in B^\varepsilon(\rho)}D(\rho'\|\sigma)
\]
[2601.10395]. In the same framework, the hypothesis-testing divergence is
\[
D_H^\varepsilon(\rho\|\sigma)\;=\;-\log\bigl[\,\min\{\Tr[\Lambda\sigma]\;:\;\Tr[\Lambda\rho]\ge1-\varepsilon,\;0\le\Lambda\le I\}\bigr]
\]
and the max-divergence is
\[
D_{\max}(\rho\|\sigma)\;=\;\inf\{\lambda:2^\lambda\,\sigma\ge\rho\}
\]
with its smoothed version obtained by minimizing \(D_{\max}\) over \(B^\varepsilon(\rho)\) [2603.09885].

A distinct construction is the induced divergence introduced from a parent quantum relative entropy \(D(\rho\|\sigma)\). For \(\varepsilon\in(0,1)\), the unnormalized induced divergence is
\[
R_D^\varepsilon(\rho\|\sigma)
\;=\;
\sup\Bigl\{\lambda\in\mathbb R\;:\;
D\!\bigl(\rho\;\big\|\;\rho+2^\lambda\,\sigma\bigr)\;\ge\;1-\varepsilon\Bigr\},
\]
and the normalized induced divergence is
\[
R^\varepsilon(\rho\|\sigma)
\;=\;
R_D^\varepsilon(\rho\|\sigma)\;+\;\bigl(\tfrac{1-\varepsilon}{\ln 2}\bigr),
\]
with the equivalent expression
\[
R^\varepsilon(\rho\|\sigma)
\;=\;
\sup\Bigl\{\lambda\;:\;
D\!\bigl(\rho\;\big\|\;(1-\varepsilon)\rho+2^\lambda\sigma\bigr)\ge0\Bigr\}
\]
[2502.13669].

The literature also distinguishes full smoothing from partial smoothing. For \(\rho_{AB}\in S_\circ(AB)\), a metric \(\Delta\), and an error \(\varepsilon\), the \(\varepsilon\)-ball with fixed \(B\)-marginal is
\[
B_{\varepsilon,\Delta}^{(\dot B)}(\rho_{AB})
=
\bigl\{\tilde\rho_{AB}\in S_\bullet(AB)\;:\;\Delta(\tilde\rho_{AB},\rho_{AB})\le\varepsilon,\;\tilde\rho_B=\rho_B\bigr\}.
\]
The corresponding partially smoothed max-relative entropy is
\[
D_{\max}^{\varepsilon,\Delta}(\dot A\!:\!B)_{\rho\|\sigma}
=
\min_{\tilde\rho_{AB}\in B_{\varepsilon,\Delta}^{(\dot A)}(\rho)}
D_{\max}\bigl(\tilde\rho_{AB}\,\big\|\,\rho_A\otimes\sigma_B\bigr),
\]
and an analogous fixed-subsystem definition exists for \(D_H\) [1807.05630].

These definitions already exhibit three non-equivalent notions of smoothing: smoothing by optimization over nearby states, smoothing through induced thresholds inside a parent divergence, and smoothing with fixed-marginal constraints. This suggests that “smoothed quantum divergence” is best understood as a family of constructions rather than a single object.

## 2. Structural principles and interpolation behavior

For induced divergences, the smoothing parameter interpolates between extremal entropy notions. If the parent divergence \(D\) is continuous in its second argument, then
\[
R^\varepsilon(\rho\|\sigma)\le D(\rho\|\sigma)-\log(1-\varepsilon),
\]
\[
\lim_{\varepsilon\to1^-}R^\varepsilon(\rho\|\sigma)=D(\rho\|\sigma),
\]
and
\[
D_{\min}(\rho\|\sigma)\le R^\varepsilon(\rho\|\sigma)\le D_{\max}(\rho\|\sigma).
\]
Moreover, if \(D_{\min}\) or \(D_{\max}\) are self-induced then
\[
\lim_{\varepsilon\to0^+}R^\varepsilon(\rho\|\sigma)=D_{\min}(\rho\|\sigma).
\]
In particular, for a sandwiched Rényi parent \(D_\alpha\) with \(\alpha\in[0,2]\), the induced divergence satisfies
\[
\lim_{\varepsilon\to0^+}R_{D_\alpha}^\varepsilon(\rho\|\sigma)=D_{\min}(\rho\|\sigma),
\quad
\lim_{\varepsilon\to1^-}R_{D_\alpha}^\varepsilon(\rho\|\sigma)=D_\alpha(\rho\|\sigma)
\]
[2502.13669].

For trace-distance smoothing, a universal structural principle was identified first in the classical setting and then lifted to the quantum setting. If \(p,q\) are probability vectors with likelihood ratios \(r_x=p_x/q_x\), then
\[
\min_{p':\tfrac12\|p'-p\|_1\le\varepsilon}D(p'\|q)=D\bigl(p^{(\varepsilon)}\|q\bigr),
\]
where \(p^{(\varepsilon)}\) is the unique “flattest” or \(\varepsilon\)-clipped approximation of \(p\) relative to \(q\),
\[
p_x^{(\varepsilon)}=q_x\max\{b,\min\{a,r_x\}\},
\]
with cutoffs \(a,b\) chosen so that the positive and negative deviations each sum to \(\varepsilon\) [2603.09885]. The same work states that, by the measured-divergence reduction,
\[
D^\varepsilon(\rho\|\sigma)=\sup_M D^\varepsilon\bigl(M(\rho)\|M(\sigma)\bigr),
\]
so the \(\varepsilon\)-clipped-vector structure underlies all smoothed divergences in the trace-distance model [2603.09885].

A complementary geometric statement appears in the analysis of Pinsker-type inequalities. Under mild axioms, the optimal convex lower bound for the smoothed divergence satisfies
\[
B_{D^\varepsilon}(T)=
\begin{cases}
0,&0\le T\le\varepsilon,\\
B_D(T-\varepsilon),&\varepsilon<T\le1,
\end{cases}
\]
so smoothing “cuts off” the region \(0\le T\le\varepsilon\) and shifts the bound to the right by \(\varepsilon\) [2601.10395]. Operationally, smoothing corresponds to allowing the input state \(\rho\) to vary in trace-distance within \(\varepsilon\), so that “almost indistinguishable” pairs can be made exactly indistinguishable [2601.10395].

## 3. Core mathematical properties

A recurring theme is that the principal smoothed divergences preserve data processing. For induced divergences, for any CPTP map or superchannel \(\Phi\),
\[
R_D^\varepsilon\bigl(\Phi(\rho)\,\big\|\,\Phi(\sigma)\bigr)\le R_D^\varepsilon(\rho\|\sigma),
\qquad
R^\varepsilon\bigl(\Phi(\rho)\|\Phi(\sigma)\bigr)\le R^\varepsilon(\rho\|\sigma)
\]
[2502.13669]. For smoothed max- and min-type divergences,
\[
D_{\max}^\varepsilon(\rho\|\sigma)\ge D_{\max}^\varepsilon(\mathcal N(\rho)\|\mathcal N(\sigma)),
\qquad
D_{\min}^\varepsilon(\rho\|\sigma)\ge D_{\min}^\varepsilon(\mathcal N(\rho)\|\mathcal N(\sigma))
\]
for any CPTP map \(\mathcal N\) [2510.02071]. Partial smoothing retains an analogous monotonicity under local CPTP maps [1807.05630].

Induced divergences also inherit scaling and Löwner monotonicity. For all \(t>0\),
\[
R_D^\varepsilon(\rho\|t\sigma)=R_D^\varepsilon(\rho\|\sigma)-\log t,
\]
and whenever \(\sigma'\succeq\sigma\),
\[
R_D^\varepsilon(\rho\|\sigma')\le R_D^\varepsilon(\rho\|\sigma)
\]
[2502.13669]. These properties are central because the induced divergence is designed to replace the hypothesis-testing divergence in position-based decoding while remaining compatible with the same monotonicity-based proof strategies [2502.13669].

Monotonicity in the smoothing parameter is explicit for trace-distance smoothed max- and min-type divergences: if \(\varepsilon_1\le\varepsilon_2\), then
\[
D_{\max}^{\varepsilon_1}(\rho\|\sigma)\ge D_{\max}^{\varepsilon_2}(\rho\|\sigma),
\qquad
D_{\min}^{\varepsilon_1}(\rho\|\sigma)\le D_{\min}^{\varepsilon_2}(\rho\|\sigma)
\]
[2510.02071]. Partial smoothing also admits continuity in \(\varepsilon\), with
\[
\bigl|D_{\max}^{\varepsilon,\Delta}-D_{\max}^{\varepsilon',\Delta}\bigr|
\le O\bigl(|\varepsilon-\varepsilon'|\bigr)
\]
with explicit logarithmic prefactors [1807.05630].

Asymptotic equipartition principles place smoothed divergences within the standard entropy hierarchy. For induced divergences built from sandwiched Rényi parents with \(\alpha\in[1,2]\),
\[
\lim_{n\to\infty}\frac1nR_{D_\alpha}^\varepsilon\!\bigl(\rho^{\otimes n}\big\|\sigma^{\otimes n}\bigr)=D(\rho\|\sigma)
\]
for fixed \(\varepsilon\in(0,1)\) [2502.13669]. For smoothed divergences between sets,
\[
\lim_{n\to\infty}\tfrac1nD_{\max}^\varepsilon(\mathcal A_n\|\mathcal B_n)
=
\lim_{n\to\infty}\tfrac1nD_{\min}^\varepsilon(\mathcal A_n\|\mathcal B_n)
=
D^\infty(\mathcal A\|\mathcal B),
\]
independent of \(\varepsilon\), where \(D^\infty\) is the regularized Umegaki-relative-entropy divergence [2510.02071].

## 4. Relations among divergence families

A major use of smoothing is to compare divergence families that are operationally natural but analytically different. For the induced divergence and \(\alpha\in[0,2]\),
\[
R_{D_\alpha}^\varepsilon(\rho\|\sigma)\ge
D_H^\varepsilon(\rho\|\sigma)-\log\frac1{1-\varepsilon},
\]
or, in normalized form,
\[
R^\varepsilon(\rho\|\sigma)\ge \underline D_H^\varepsilon(\rho\|\sigma)
\]
[2502.13669]. The same work also states explicit relations with information-spectrum smoothing \(\underline D_s^\varepsilon\) and \(\tilde D_{\max}^\varepsilon\), and gives the sandwiched Rényi representation
\[
R_{D_\alpha}^\varepsilon(\rho\|\sigma)
=
\sup\Bigl\{\lambda:\;Q_\alpha\bigl(\rho\,\big\|\rho+2^\lambda\sigma\bigr)\ge (1-\varepsilon)^{\alpha-1}\Bigr\}
\]
for \(\alpha\in(1,\infty)\) [2502.13669].

Recent universal bounds place smoothed Rényi divergences of one order against unsmoothed Rényi divergences of another order. For \(\varepsilon\in(0,1)\), fixed \(\alpha,\beta\), and optimal state-independent constants \(\mu(\varepsilon,\alpha,\beta)\) and \(\nu(\varepsilon,\alpha,\beta)\), the bounds take the form
\[
D_\beta^\varepsilon(\rho\|\sigma)\le D_\alpha(\rho\|\sigma)+\Delta(\varepsilon,\alpha,\beta)
\]
in the regimes \(0<\alpha<\beta<1\) and \(\beta>\alpha>1\), with \(\Delta\) given in closed form, and
\[
D_\beta^\varepsilon(\rho\|\sigma)\ge
D_\alpha(\rho\|\sigma)-(1-\varepsilon)^{-1}\Bigl(\tfrac1{\beta-1}+\tfrac1{1-\alpha}\Bigr)
\]
whenever \(\beta>1>\alpha\) [2603.09885]. The same paper states that the correction terms are optimal among all universal, state-independent inequalities of this type [2603.09885].

For the hypothesis-testing divergence, the same universal comparison yields, for \(\alpha>1\),
\[
D_H^\varepsilon(\rho\|\sigma)\le
D_\alpha(\rho\|\sigma)+\frac{\alpha}{\alpha-1}\log\frac1{1-\varepsilon},
\]
and this correction is stated to be optimal [2603.09885]. For \(\alpha\in(0,1)\), a piecewise lower bound is given in terms of \(\varepsilon\) and \(\alpha\) [2603.09885]. This places \(D_H^\varepsilon\) within the same order-comparison framework as the Rényi family.

Pinsker-type inequalities provide another comparison mechanism by replacing divergence with trace distance. For example,
\[
B_{D_2}(T)=
\begin{cases}
\log(1+4T^2),&0\le T\le\tfrac12,\\
\log[1/(1-T)],&\tfrac12\le T\le1,
\end{cases}
\]
\[
B_{D_\infty}(T)=\log[1/(1-T)],
\qquad
B_{D_{1/2}}(T)=\log[1/(1-T^2)]
\]
for the collision, max, and fidelity divergences respectively [2601.10395]. After smoothing, the lower bound is shifted by \(\varepsilon\), which gives an explicit bridge from difficult finite-size divergences to the trace distance [2601.10395].

## 5. Partial smoothing and fixed-marginal constraints

Partial smoothing was introduced to address the mismatch between standard smoothing and operational tasks in which some subsystems are not meant to vary. In the partially smoothed framework, smoothing is performed only over nearby states that preserve a specified marginal exactly, or in a variant formulation satisfy an upper-bound constraint on that marginal [1807.05630], [1905.08268].

For bipartite quantities, one defines fixed-subsystem versions of \(D_{\max}\) and \(D_H\), such as
\[
D_{\max}^{\varepsilon,\Delta}(\dot A\!:\!B)_{\rho\|\sigma}
=
\min_{\tilde\rho_{AB}\in B_{\varepsilon,\Delta}^{(\dot A)}(\rho)}
D_{\max}\bigl(\tilde\rho_{AB}\|\rho_A\otimes\sigma_B\bigr),
\]
and similarly for the hypothesis-testing divergence [1807.05630]. These measures obey data processing, monotonicity under partial trace, additivity up to \(\varepsilon\)-splitting, and chain-rule type bounds in the partially constrained setting [1807.05630].

The distinction between global and partial smoothing is especially visible for conditional min-entropy. The globally smoothed quantity is identified with a smoothed max-divergence against \(I_A\otimes\rho_B\), whereas the partially smoothed version imposes the additional constraint \(\Tr_A\sigma_{AB}\le\rho_B\) [1905.08268]. For i.i.d. pure states \(\rho_{AR}^{\otimes n}\), the partially smoothed conditional min-entropy satisfies
\[
H_{\min}^{\varepsilon,P}(A^n\,|\,\!\cdot R^n)_{\rho^{\otimes n}}
=
nH(A)_\rho+\sqrt{nV(A)_\rho}\,\Phi^{-1}\!\bigl(\sqrt{1-\varepsilon^2}\bigr)+O(\log n),
\]
whereas global smoothing yields a second-order coefficient \(\Phi^{-1}(1-\varepsilon^2)\) [1905.08268]. The paper states that for general pure states the second-order term differs, while for several natural classes of states partial and global smoothing coincide [1905.08268].

This fixed-marginal approach is operationally useful because unconstrained smoothing can introduce penalties not present in the task formulation. The literature explicitly states that in classical state-splitting unconstrained smoothing forces an extra \(\log|X|\)-penalty, while partial smoothing removes it, and that partially smoothed quantities yield tight second-order behavior in several classical side-information settings [1807.05630].

## 6. Operational roles in one-shot and asymptotic information theory

Smoothed divergences are primarily motivated by one-shot coding theorems. In classical communication over a quantum channel \(N:X\to B\) with classical input, the induced collision-mutual information is defined as
\[
R I_2^\varepsilon(X:B)_N
=
\sup_{p(x)}R_{D_2}^\varepsilon\!\Bigl(\sigma_p^{XB}\,\Big\|\,\sigma_p^X\otimes\sigma_p^B\Bigr),
\]
where \(\sigma_p^{XB}=\sum_x p(x)|x\rangle\langle x|\otimes N(|x\rangle\langle x|)\). The one-shot achievable lower bound is
\[
C_\varepsilon^{(1)}(N)\ge R I_2^\varepsilon(X:B)_N+\log(1-\varepsilon),
\]
which is stated to recover and strengthen previous bounds based on hypothesis-testing divergences [2502.13669].

For entanglement-assisted quantum state redistribution with pure source \(\rho^{RAA'B}\), the paper defines a single-shot smoothed conditional mutual information
\[
R I_2^\varepsilon(A':R\mid B)_\rho
=
I_2^{\delta_0}(RB\!:\!A')_\rho - R I_2^{\delta_1}(B\!:\!A')_\rho,
\]
and gives the one-shot cost bound
\[
Q_\varepsilon(\rho^{AA'B})
\le
\tfrac12 R I_2^\varepsilon(A':R\mid B)_\rho+\tfrac12\log\frac1{\delta'}
\]
for suitably small \(\delta_0,\delta_1,\delta'\) satisfying
\[
\delta'-\sqrt{2(\delta_0+\delta_1)}-\sqrt{2\delta_0}>0
\]
[2502.13669]. The stated refinement is that earlier state-redistribution bounds using \(D_{\max}\) can be sharpened by replacing it with \(D_2\) and induced smoothing [2502.13669].

Partial smoothing also has direct one-shot operational meanings. For classical state splitting,
\[
I_{\max}^{\varepsilon,T}(\dot X:Y)_P
\le R \le
I_{\max}^{\varepsilon-\delta,T}(\dot X:Y)_P+\log\log\frac1\delta+1
\]
for \(0<\delta<\varepsilon\), and the i.i.d. limit gives the second-order expansion
\[
R/n=I(X;Y)+\sqrt{V(X;Y)/n}\,\Phi^{-1}(\varepsilon)+O\!\bigl(\tfrac{\log n}{n}\bigr)
\]
[1807.05630]. For privacy amplification,
\[
H_{\min}^{-\delta,P}(X|\dot B)-4\log\frac1\delta
\le \ell \le
H_{\min}^{\varepsilon,P}(X|\dot B),
\]
and in the classical side-information case the second-order expansion is given in terms of \(H(X|Y)\) and \(V(X|Y)\) [1807.05630]. For state merging,
\[
-\,H_{\min}^{\varepsilon,P}(A|\dot R)_\rho \le E
\le -\,H_{\min}^{-\delta,P}(A|\dot R)_\rho+4\log\frac1\delta
\]
for entanglement cost, while the classical communication cost satisfies
\[
I_{\max}^{\varepsilon,P}(\dot R:A)_\rho \le C
\le I_{\max}^{-\delta,P}(\dot R:A)_\rho+4\log\frac1\delta
\]
[1807.05630].

The second-order asymptotics of quantum compression provide a further application. For source state \(\rho_A\), blocklength \(n\), and entanglement-fidelity error \(\delta\), the minimal compression size obeys
\[
M^*(\rho_A^{\otimes n},\delta)
=
nH(A)_\rho+\sqrt{nV(A)_\rho}\,\Phi^{-1}(\sqrt{1-\delta})+O(\log n),
\]
and the straightforward protocol of cutting off the eigenspace of least weight is stated to be asymptotically optimal at second order [1905.08268].

Smoothed divergences between sets also acquire an exact operational meaning in the resource theory of asymmetric distinguishability with partial information. For compact convex sets \(\mathcal A,\mathcal B\),
\[
D_{\rm distill}^\varepsilon(\mathcal A,\mathcal B)=D_{\min}^\varepsilon(\mathcal A\|\mathcal B),
\qquad
D_{\rm dilute}^\varepsilon(\mathcal A,\mathcal B)=D_{\max}^\varepsilon(\mathcal A\|\mathcal B)
\]
[2510.02071]. Under regularity assumptions, both asymptotic rates converge to the same regularized Umegaki-relative-entropy divergence \(D^\infty(\mathcal A\|\mathcal B)\), and the optimal asymptotic conversion rate between two resource objects is the ratio of the corresponding regularized divergences [2510.02071].

## 7. Conceptual distinctions, common misconceptions, and current directions

A common misconception is that smoothing merely introduces a technical \(\varepsilon\)-slack into an otherwise unchanged divergence. The recent literature instead presents several inequivalent smoothing mechanisms: infimum-based trace-distance smoothing [2601.10395], induced smoothing through the condition \(D(\rho\|\rho+2^\lambda\sigma)\ge1-\varepsilon\) [2502.13669], and partial smoothing with fixed marginals [1807.05630]. These choices lead to different interpolation limits, distinct second-order terms, and different one-shot achievability statements.

Another misconception is that all smoothed divergences are interchangeable in finite-blocklength analysis. The papers summarized here show that this is not generally the case. Induced divergences are designed to replace the hypothesis-testing divergence in position-based decoding and can yield tighter achievability bounds [2502.13669]. Partial smoothing can remove artifacts of global smoothing in tasks with fixed side information [1807.05630]. Universal order-comparison bounds demonstrate that some relations between smoothed and unsmoothed divergences are sharp, while outside specified parameter regimes no nontrivial universal bound exists [2603.09885].

A further point of clarification concerns geometry. In the trace-distance model, the optimizer of the classical smoothing problem is a clipped probability vector, and the same geometry underlies the quantum case through measured-divergence reduction [2603.09885]. By contrast, induced divergences are not defined by minimizing over a trace-distance ball; their geometry is encoded in a threshold condition on \(\rho+2^\lambda\sigma\) relative to a parent divergence [2502.13669].

Several research directions are explicitly identified in the cited works. These include quantum second-order expansions for partially smoothed quantities in general, extensions of partial smoothing to Rényi divergences and smooth Rényi entropies with fixed marginals, multipartite partial smoothing, and applications to one-shot quantum channel coding and quantum joint typicality conjectures [1807.05630]. A plausible implication is that future progress will depend on combining the geometric understanding of clipped-vector smoothing with the task-specific advantages of structured and induced smoothing.

Source: https://www.emergentmind.com/topics/smoothed-quantum-divergences