---
title: Smoothed Classical Divergences
url: https://www.emergentmind.com/topics/smoothed-classical-divergences
type: topic
---

# Smoothed Classical Divergences

In the literature considered here, smoothed classical divergences comprise several distinct regularization mechanisms for classical discrepancy functionals. One line of work identifies a canonical divergence on flat statistical manifolds and shows that, on the manifold of positive measures, it coincides exactly with the classical $\alpha$-divergence [1907.11122]. Another defines smoothing by minimization over a total-variation ball, $D^\varepsilon(p\|q):=\min_{p'\in B^\varepsilon(p)}D(p'\|q)$, and proves that the optimizer is a divergence-independent clipped probability vector [2603.09885]. A further line regularizes Rényi divergence by infimal convolution with an integral probability metric, thereby interpolating between Rényi divergences and IPMs [2210.04974]. In statistical inference for continuous models, smoothing can also mean smoothing both the data density and the model density before divergence minimization, as in the Basu–Lindsay approach for minimum $S$-divergence estimation [1408.1239].

## 1. Formal setting and recurrent constructions

A recurring framework starts from a classical divergence
\[
D:\bigcup_{d\in\mathbb N}\big(\Delta_d\times \Delta_d\big)\to \mathbb R\cup\{\infty\},
\]
assumed to satisfy the data processing inequality
\[
D(Ep\|Eq)\le D(p\|q)
\qquad\text{for every column-stochastic map }E.
\]
In the total-variation formulation, smoothing is defined through the ball
\[
B^\varepsilon(p):=\left\{p'\in\Delta_d:\frac12\|p-p'\|_1\le \varepsilon\right\},
\]
and the smoothed divergence is
\[
D^\varepsilon(p\|q):=\min_{p'\in B^\varepsilon(p)}D(p'\|q).
\]
The same work emphasizes that $D^\varepsilon$ is again a classical divergence, because both the feasible set $B^\varepsilon(p)$ and $D$ are compatible with stochastic maps [2603.09885].

A different regularization pattern is infimal convolution. For a test-function space $\Gamma\subset C(X)$, the infimal-convolution $\Gamma$-Rényi divergence is defined by
\[
R_\alpha^{\Gamma,IC}(P\|Q)\coloneqq \inf_{\eta\in\mathcal{P}(X)}\{R_\alpha(P\|\eta)+W^\Gamma(Q,\eta)\},
\]
where
\[
W^\Gamma(\mu,\nu)\coloneqq \sup_{g\in\Gamma}\left\{\int gd\mu-\int gd\nu\right\}.
\]
Here smoothing acts on the second argument through an auxiliary measure $\eta$ and a geometric penalty determined by $\Gamma$ [2210.04974].

In continuous-model inference, smoothing can instead be applied to the compared objects themselves. The Basu–Lindsay method replaces the empirical distribution by a kernel density estimator
\[
g_n^*(x)=\frac1n\sum_{i=1}^n W(x,X_i,h_n),
\]
and compares it not with the raw model density $f_\theta$, but with the smoothed model
\[
f_\theta^*(x)=\int W(x,y,h)\,dF_\theta(y).
\]
The resulting estimator minimizes $S_{(\alpha,\lambda)}(g_n^*,f_\theta^*)$ rather than the unsmoothed divergence [1408.1239].

Taken together, these constructions indicate that smoothing need not refer to a single operation. It may act on the feasible set, on one argument of the divergence, on both arguments, or on the underlying geometric structure.

## 2. Information-geometric smoothing and the classical $\alpha$-divergence

The information-geometric formulation begins with a smooth manifold $M$ equipped with a dualistic structure
\[
(M,g,\nabla,\nabla^*),
\]
where $g$ is a Riemannian metric and $\nabla,\nabla^*$ are affine connections dual with respect to $g$ in the sense that
\[
X\, g(Y,Z)= g(\nabla_X Y,Z)+g(Y,\nabla^*_X Z).
\]
If both connections are torsion-free, the manifold is statistical; if both have zero curvature,
\[
R(\nabla)=0,\qquad R(\nabla^*)=0,
\]
the structure is flat [1907.11122].

On the manifold of positive measures, the $\alpha$-connections interpolate between the mixture and exponential connections:
\[
\nabla^{(\alpha)}=\frac{1-\alpha}{2}\,\nabla^{(m)}+\frac{1+\alpha}{2}\,\nabla^{(e)}.
\]
These $\alpha$-connections are dual with respect to the Fisher metric, and the pair $(g_F,\nabla^{(\alpha)},\nabla^{(-\alpha)})$ is dually flat [1907.11122].

The canonical divergence introduced by Ay and Amari is defined by integrating the squared norm of the $\nabla$-geodesic velocity:
\[
D(p,q)=\int_0^1 \left\lVert \dot\gamma(t)\right\rVert^2\,dt,
\]
where $\gamma:[0,1]\to M$ is the $\nabla$-geodesic from $q$ to $p$. In the general construction, the inverse exponential map is used to define the initial velocity and to transport it along the geodesic. On dually flat manifolds, this canonical divergence specializes to a Bregman-type canonical divergence; in the self-dual Levi-Civita case, it becomes $\frac12 d^2$ [1907.11122].

For the manifold
\[
M^+=\left\{p=\sum_{i=1}^n p_i\, d_i \in \mathbb{R}^n \ \middle|\ p_i>0\ \forall i\right\},
\]
equipped with the Fisher metric
\[
g_p(X,Y)=\sum_{i=1}^n \frac{X_iY_i}{p_i},
\]
the paper computes the canonical divergence explicitly and proves
\[
D(p,q)=D^{(\alpha)}(p,q).
\]
That is, the canonical divergence coincides exactly with the classical $\alpha$-divergence on positive measures [1907.11122].

The same analysis records the limiting behavior of the classical family. As $\alpha\to -1$, the divergence converges to the Kullback–Leibler divergence,
\[
\lim_{\alpha\to -1} D^{(\alpha)}(p,q) = \sum_{i=1}^n p_i \log\frac{p_i}{q_i},
\]
while $\alpha\to +1$ yields the reverse KL form. The paper also notes that for $\alpha=0$ the divergence is closely related to the Hellinger-type geometry, and that the parameter change $\alpha=1-2q$ connects the family to Tsallis/$q$-divergence [1907.11122]. This suggests that, in flat $\alpha$-geometry, smoothing is realized intrinsically rather than by external mollification.

## 3. Total-variation smoothing and the clipping principle

For classical divergences satisfying data processing, the central structural result is that smoothing over a total-variation ball has a divergence-independent optimizer. Let
\[
r_x:=\frac{p_x}{q_x},
\]
with the likelihood ratios ordered as $r_1\ge r_2\ge \cdots \ge r_d$. Define the clipping thresholds
\[
a:=\max_{m\in[d]}\frac{\sum_{x\in[m]}p_x-\varepsilon}{\sum_{x\in[m]}q_x},
\qquad
b:=\min_{\ell\in[d]}\frac{\sum_{x=\ell}^d p_x+\varepsilon}{\sum_{x=\ell}^d q_x}.
\]
The clipped vector $p^{(\varepsilon)}$ is then
\[
p_x^{(\varepsilon)} = q_x\,\max\{b,\min\{a,r_x\}\}
\qquad\forall x\in[d].
\]
Equivalently,
\[
p_x^{(\varepsilon)}= \begin{cases}
a\,q_x,& r_x>a,\\[2mm]
p_x,& b\le r_x\le a,\\[2mm]
b\,q_x,& r_x<b.
\end{cases}
\]
This is the unique vector obtained by clipping the likelihood ratio $r_x$ to the interval $[b,a]$ [2603.09885].

The main theorem states that for every classical divergence $D$,
\[
D^\varepsilon(p\|q)=D\big(p^{(\varepsilon)}\|q\big).
\]
Hence the minimizer in the smoothing problem is always the same clipped vector $p^{(\varepsilon)}$, independent of the specific divergence. The divergence dependence enters only through the final evaluation at the clipped point [2603.09885].

The geometric mechanism is majorization. The total-variation ball $B^\varepsilon(p)$ has extremal elements under majorization, and for minimization the relevant one is the flattest approximation. Operationally, smoothing allows redistribution of at most $\varepsilon$ total mass; the optimizer decreases overly large coordinates, increases overly small coordinates, and leaves intermediate coordinates unchanged. In relative-majorization language, the clipped construction is extremal in the $\varepsilon$-ball, which is the order-theoretic reason for divergence-independence [2603.09885].

A significant technical reduction uses relative majorization and uniform reference. For rational
\[
q=\left(\frac{k_1}{k},\dots,\frac{k_d}{k}\right),
\]
the pair $(p,q)$ can be converted into a pair $(t,u)$ with $u$ uniform, and for every classical divergence $D$,
\[
D(p\|q)=D(t\|u).
\]
The uniform-reference case is then lifted back to general $q$ [2603.09885].

## 4. Function-space regularized Rényi divergences

The function-space regularized Rényi divergences place Rényi divergence into an explicit smoothed-divergence framework. For $\alpha\in(0,1)\cup(1,\infty)$, the classical Rényi divergence between $P$ and $Q$ is paired with an IPM penalty through the infimal convolution
\[
R_\alpha^{\Gamma,IC}(P\|Q)
=\inf_{\eta\in\mathcal P(X)}
\{R_\alpha(P\|\eta)+W^\Gamma(Q,\eta)\}.
\]
The function space $\Gamma$ specifies the regularization geometry: choosing $\Gamma=Lip^1(X)$ gives Wasserstein regularization, bounded functions give total variation, $Lip^1\cap L^\infty$ gives Dudley-type metrics, and RKHS unit balls give MMD-type regularization [2210.04974].

The computational core is the dual representation obtained by Fenchel–Rockafellar duality:
\[
R_\alpha^{\Gamma,IC}(P\|Q)=\sup_{g\in \Gamma:g<0}\left\{\int gdQ +\frac{1}{\alpha-1}\log\int |g|^{(\alpha-1)/\alpha} dP\right\}+\alpha^{-1}(\log\alpha +1).
\]
The paper emphasizes that this formula avoids the risk-sensitive exponential terms appearing in the Donsker–Varadhan representation and therefore exhibits lower variance, making it well-behaved when $\alpha>1$ [2210.04974].

Theorem 1 establishes the basic interpolation property:
\[
R_\alpha^{\Gamma,IC}(P\|Q)\leq \min\{R_\alpha(P\|Q),W^\Gamma(Q,P)\},
\]
together with nonnegativity and equality at $P=Q$. If $\Gamma$ is strictly admissible, the divergence property also holds. The same theorem proves convexity in $Q$, joint convexity in $(P,Q)$ when $\alpha\in(0,1)$, and lower semicontinuity [2210.04974].

The limiting regimes make the interpolation precise. For admissible $\Gamma$,
\[
\lim_{\delta\to0^+}\frac{1}{\delta}R_\alpha^{\delta\Gamma,IC}(P\|Q)=W^\Gamma(Q,P),
\]
while for strictly admissible $\Gamma$,
\[
\lim_{L\to\infty}R_\alpha^{L\Gamma,IC}(P\|Q)=R_\alpha(P\|Q).
\]
At $\alpha\to1^-$, the construction yields the expected reverse-KL-type regularization; at $\alpha\to\infty$, the rescaled family converges to a regularized worst-case-regret divergence [2210.04974].

A major consequence is the removal of the absolute-continuity obstruction. Classical $R_\alpha(P\|Q)$ is infinite for $\alpha>1$ unless $P\ll Q$, whereas the infimal-convolution regularized versions can compare measures that are not mutually absolutely continuous, including empirical distributions and low-dimensional supports [2210.04974]. The paper also records a data-processing inequality for the regularized divergences. By contrast, naive regularizations based on simply restricting the Donsker–Varadhan test-function space can fail to satisfy $R_\alpha^{\Gamma,DV}\le W^\Gamma$ for $\alpha>1$ and can be numerically unstable [2210.04974].

## 5. Binary reduction and Pinsker-type bounds for smoothed divergences

For a broad family of data-processing divergences, optimal lower bounds in terms of trace or variational distance reduce to binary classical states. Writing
\[
T(r,s)=|r-s|
\]
for binary distributions, the fundamental linear lower-bound problem is
\[
L_D(\lambda)=\inf_{\rho,\sigma}\{D(\rho\|\sigma)-\lambda T(\rho,\sigma)\},
\]
and Theorem 1 gives the binary reduction
\[
L_{D}(\lambda) =\inf_{0\le s\le r\le 1}\Big\{D_{\mathrm{bin}}(r\|s)-\lambda(r-s)\Big\}.
\]
The optimal convex bound
\[
D(\rho\|\sigma)\ge B_D(T(\rho,\sigma))
\]
is also attained on two-dimensional classical states [2601.10395].

For smoothed divergences,
\[
D^\varepsilon(\rho\|\sigma)=\inf_{\rho'\in B^\varepsilon(\rho)} D(\rho'\|\sigma),
\]
the paper proves a general shift-and-cutoff theorem. Assuming data processing, non-negativity, convexity in the first argument, and faithfulness, the convex lower bound becomes
\[
B_{D^\varepsilon}(T)= \begin{cases}
0, & 0\le T\le \varepsilon,\\[1mm]
B_D(T-\varepsilon), & \varepsilon<T\le 1.
\end{cases}
\]
Thus the unsmoothed bound is shifted to the right by $\varepsilon$ and set to zero on $[0,\varepsilon]$ [2601.10395].

The paper provides explicit bounds for several classical and quantum divergences whose binary optimization is classical in form.

| Divergence | Lower bound in terms of $T$ | Smoothed form |
|---|---|---|
| Umegaki divergence | $\frac{2}{\ln 2}\,T^2$ as the simple Pinsker form | Shift rule applies |
| Fidelity divergence | $\log\!\left(\frac{1}{1-T^2}\right)$ | Shift rule applies |
| Neyman/Pearson $\chi^2$ | $4T^2$ for $0\le T\le \frac12$, then $\frac{T}{1-T}$ | Shift rule applies |
| Max divergence | $\log\!\left(\frac{1}{1-T}\right)$ | For $D_\infty^\varepsilon$, zero on $[0,\varepsilon]$, then $\log\!\left(\frac{1}{1-(T-\varepsilon)}\right)$ |

The same work notes several structural features: the bounds are often piecewise and meet continuously with matching derivatives; standard Pinsker-type estimates are loose for large $T$; for Rényi divergences with $\alpha\in[1,\infty]$, the large-$T$ branch $\log(1/(1-T))$ is universal; and smoothing can remove the divergence of unsmoothed bounds as $T\to1$ [2601.10395].

## 6. Statistical inference, smooth surrogates, and constrained optimization

In continuous parametric models, the minimum $S$-divergence estimator requires smoothing because the empirical distribution is discrete while the model density is continuous. The $S$-divergence family is indexed by $\alpha\ge0$ and $\lambda\in\mathbb R$,
\[
S_{(\alpha,\lambda)}(g,f) = \frac{1}{A}\int f^{1+\alpha} -\frac{1+\alpha}{AB}\int f^B g^A +\frac{1}{B}\int g^{1+\alpha},
\]
with
\[
A = 1+\lambda(1-\alpha), \qquad B = \alpha - \lambda(1-\alpha),
\]
and includes the Cressie–Read power divergence at $\alpha=0$, the density power divergence at $\lambda=0$, and the $L_2$ divergence at $\alpha=1$ [1408.1239].

Using the Basu–Lindsay approach, one smooths both the data and the model with the same kernel and minimizes
\[
S_{(\alpha,\lambda)}(g_n^*, f_\theta^*).
\]
The resulting minimum $S^*$-divergence estimator satisfies consistency and asymptotic normality under identifiability, common-support, smoothness, domination, and positive-definiteness conditions. At the model, both the influence function and the asymptotic distribution are independent of $\lambda$, while second-order influence analysis shows explicitly that $\lambda$ enters at second order [1408.1239]. This corrects the common first-order impression that the choice of $\lambda$ is asymptotically irrelevant.

A different smoothing goal is the replacement of nonsmooth $\ell_1$ discrepancies by smooth divergence surrogates. The smooth generator $\varphi_{\alpha,\beta,\widetilde c}$ introduced for generalized $\varphi$-divergences is strictly convex and $C^\infty$, satisfies
\[
\varphi_{\alpha,\beta,\widetilde c}(1)=0,
\]
and converges pointwise as
\[
\lim_{\alpha \to 0_+}\varphi_{\alpha,\beta,\widetilde c}(t) = \widetilde{c}\,\beta\,|t-1|.
\]
Consequently,
\[
\lim_{\alpha\to 0_+} D_{\varphi_{\alpha,\beta,\widetilde c}}(\mathbf{Q},\mathbf{P})
=
\widetilde{c}\,\beta\,\|\mathbf{Q}-\mathbf{P}\|_1,
\]
and the scaled shift divergence similarly converges to a weighted $\ell_1$-distance. The paper also states the upper bound
\[
\varphi_{\alpha,\beta,\widetilde c}(t)\leq \widetilde{c}\,\beta\,|t-1|,
\]
with equality iff $t=1$ [2511.00219]. Here smoothing is a smooth approximation of a nonsmooth classical target.

In inverse problems, smoothing also appears through regularization and scale invariance. The generic objective is
\[
J(x)=D_1(y\|m(x))+\gamma D_2(x\|x_d),
\]
with $D_2$ interpreted as a Tikhonov-type regularizer and $x_d$ as a default or smoothed solution. To handle nonnegative vectors with arbitrary total mass, the paper introduces scale-invariant divergences satisfying
\[
DI(p\|q)=DI(p\|a q),\qquad a>0,
\]
often through an invariance factor
\[
K_0(p,q)=\arg\min_{K>0}D(p\|Kq).
\]
A central warning is that simplified probability-density forms of classical divergences can be unsuitable for inverse-problem optimization because the gradient with respect to the model argument may fail to vanish at $p=q$; the KL example is used to illustrate this point [2003.01411]. This is a distinct but related sense in which smoothing modifies classical divergences so that they remain compatible with optimization, constraints, and regularization.

These developments collectively show that smoothed classical divergences are not a single family but a class of constructions unified by a common purpose: preserving the discriminative content of classical divergences while modifying geometry, admissible perturbations, arguments, or generators so that the resulting objects remain tractable in information geometry, sharp inequalities, robust inference, and constrained optimization.

Source: https://www.emergentmind.com/topics/smoothed-classical-divergences