---
title: Channel Rényi Information
url: https://www.emergentmind.com/topics/channel-renyi-information
type: topic
---

# Channel Rényi Information

Searching arXiv for recent and foundational papers on channel Rényi information and closely related Rényi mutual information/capacity results.
Channel Rényi information is the family of channel-information quantities obtained by replacing Shannon entropy or Kullback–Leibler divergence with Rényi entropies and Rényi divergences of order \(\alpha\). In the classical setting this includes Rényi capacity, minimax redundancy, Sibson mutual information, and Augustin-type quantities; in the quantum setting it includes Petz, sandwiched, log-Euclidean, \(\alpha\)-\(z\), and geometric variants built from the corresponding quantum Rényi divergences. The subject interpolates between support-sensitive behavior at \(\alpha=0\), the Shannon/KL case at \(\alpha=1\), and worst-case regret at \(\alpha=\infty\), and it is used in channel coding, leakage analysis, hypothesis testing, resolvability, and error-exponent theory [1206.2459][1811.04218].

## 1. Classical definitions and the capacity–redundancy framework

A standard starting point is the Rényi divergence
\[
D_{\alpha}(P\Vert Q)=\frac{1}{\alpha-1}\ln\int p^\alpha q^{1-\alpha}\,d\mu,
\]
or, on finite alphabets,
\[
D_{\alpha}(P\Vert Q)=\frac{1}{\alpha-1}\ln\sum_i p_i^\alpha q_i^{1-\alpha}.
\]
Its extended orders are
\[
D_0(P\Vert Q)=-\ln Q(p>0),\qquad
D_1(P\Vert Q)=D(P\Vert Q),\qquad
D_\infty(P\Vert Q)=\ln \esssup_P \frac{p}{q}.
\]
Thus the order-\(1\) case is exactly the Kullback–Leibler divergence, while \(\alpha=0\) and \(\alpha=\infty\) capture support-based and worst-case behavior, respectively [1206.2459].

For a family of channel output laws \(\{P_\theta\mid \theta\in\Theta\}\), the classical Rényi capacity and minimax redundancy are defined by
\[
C_\alpha=\sup_\pi \inf_Q \int D_\alpha(P_\theta\Vert Q)\,\pi(d\theta),\qquad
R_\alpha=\inf_Q \sup_\theta D_\alpha(P_\theta\Vert Q).
\]
A central result is that
\[
C_\alpha=R_\alpha\qquad \text{for all } \alpha\in[0,\infty]
\]
when the output alphabet is finite, and this is extended to general input spaces \(\Theta\) as well. For finite alphabets, \(C_\alpha\) is the natural Rényi analogue of Shannon capacity, and \(R_\alpha\) is the matching minimax redundancy. At \(\alpha=\infty\), the minimax redundancy becomes worst-case regret, and the unique minimizer is the Shtarkov or normalized maximum-likelihood distribution \(S\) whenever \(R_\infty<\infty\) [1206.2459].

The same framework is supported by a generalized Pythagorean inequality on \(\alpha\)-convex sets, monotonicity of \(D_\alpha(P\Vert Q)\) in \(\alpha\), convexity in the reference distribution \(Q\) for all \(\alpha\in[0,\infty]\), and continuity and compactness properties sufficient for minimax arguments. These analytic properties are not incidental: they are part of what makes Rényi channel quantities behave as genuine information measures rather than merely formal entropy substitutions [1206.2459].

## 2. Sibson mutual information, leakage, and event-wise interpretations

For a discrete memoryless channel \(P_{Y|X}\), the order-\(\alpha\) Sibson mutual information is
\[
I_{\alpha}^{\mathrm S}(P_X;P_{Y|X})
=
\frac{\alpha}{\alpha-1}
\log
\sum_{y\in Y}
\left(
\sum_{x\in X}P_X(x)P_{Y|X}^{\alpha}(y|x)
\right)^{1/\alpha},
\]
and the corresponding Rényi capacity is
\[
C_\alpha(P_{Y|X})=\max_{P_X} I_\alpha^{\mathrm S}(P_X;P_{Y|X}).
\]
This is one of the principal classical channel Rényi-information quantities, and recent work interprets it directly as a leakage functional rather than only as a divergence-based analogue of Shannon mutual information [2405.00423][2510.06622].

A key identity rewrites Sibson mutual information as
\[
I_{\alpha}^{\mathrm S}(P_X;P_{Y|X})
=
\frac{\alpha}{\alpha-1}
\log
E_{P_Y}
\!\left[
\exp\!\left(
\frac{\alpha-1}{\alpha}D_\alpha(P_{X|Y=y}\Vert P_X)
\right)
\right].
\]
With
\[
\tilde f(t)=\exp\!\left(\frac{\alpha-1}{\alpha}t\right),
\]
this means that Sibson mutual information is the \(\tilde f\)-mean of the output-wise divergences \(D_\alpha(P_{X|Y=y}\Vert P_X)\). The same papers show that, for each output symbol \(y\),
\[
D_\alpha(P_{X|Y=y}\Vert P_X)
=
\max_{P_{\hat X|Y=y}}
\tilde D_\alpha\!\left(P_{\hat X|Y=y}\Vert P_X \mid P_{X|Y=y}\right),
\]
so the posterior-to-prior Rényi divergence is the maximum achievable \(\tilde f\)-mean information gain at that output event. In this sense, Sibson mutual information aggregates optimal elementary leakages over the channel output alphabet [2405.00423][2510.06622].

The same leakage viewpoint yields a \(Y\)-elementary \(\alpha\)-leakage
\[
L_\alpha(U\to y)
=
H_\alpha(P_U)-H_\alpha(P_{U|Y=y})
=
\frac{1}{\alpha-1}
\log
\frac{\sum_u P_{U|Y}^\alpha(u|y)}{\sum_u P_U^\alpha(u)},
\]
defined for all \(\alpha\in[0,\infty)\). Maximizing this over all attributes \(U\) of the channel input \(X\) gives
\[
\sup_{P_{U|X}} L_\alpha(U\to y)=D_\alpha(P_{X|Y=y}\Vert P_X),
\]
and the \(\alpha=\infty\) case recovers the pointwise-maximal-leakage-style quantity. A plausible implication is that the Sibson-capacity viewpoint and the leakage viewpoint are not competing formalisms but two representations of the same order-\(\alpha\) channel-information structure [2510.06622].

## 3. Operational meanings: hypothesis testing, resolvability, and secrecy

An operational interpretation of Rényi mutual information is given by composite hypothesis testing. For a bipartite state \(\rho_{AB}\) and a reference state \(\tau_A\), the relevant quantities are
\[
I_{\alpha}(\rho_{AB}\|\tau_A)=\inf_{\sigma_B}D_{\alpha}(\rho_{AB}\|\tau_A\otimes \sigma_B),
\qquad
\widetilde I_{\alpha}(\rho_{AB}\|\tau_A)=\inf_{\sigma_B}\widetilde D_{\alpha}(\rho_{AB}\|\tau_A\otimes \sigma_B).
\]
The null hypothesis is \(\rho_{AB}^{\otimes n}\), while the alternative consists of all product states \(\tau_A^{\otimes n}\otimes \sigma_{B^n}\). In the Hoeffding regime,
\[
\lim_{n\to\infty}
\left\{
-\frac1n
\log
\hat\alpha\!\left(e^{-nR};\rho_{AB}^{\otimes n}\Big\|\tau_A^{\otimes n}\right)
\right\}
=
\sup_{s\in(0,1)}
\left\{
\frac{1-s}{s}\big(I_s(\rho_{AB}\|\tau_A)-R\big)
\right\},
\]
whereas in the strong-converse regime,
\[
\lim_{n\to\infty}
\left\{
-\frac1n
\log\!\left(
1-\hat\alpha\!\left(e^{-nR};\rho_{AB}^{\otimes n}\Big\|\tau_A^{\otimes n}\right)
\right)
\right\}
=
\sup_{s>1}
\left\{
\frac{s-1}{s}\big(R-\widetilde I_s(\rho_{AB}\|\tau_A)\big)
\right\}.
\]
In the classical specialization, these formulas provide an operational interpretation of classical Rényi mutual information and conditional entropy as optimal asymmetric-testing exponents against decoupled alternatives [1408.6894].

Rényi channel information also enters through resolvability. Given a target output law \(Q_Y\), Rényi resolvability is the minimum rate \(R_{1+s}(P_{Y|X},Q_Y)\) required for a channel input process to make the output distribution approximate \(Q_Y^n\) in Rényi divergence. The normalized and unnormalized Rényi resolvabilities coincide. For \(s\in(-1,0]\), the resolvability equals
\[
\min_{P_X\in\mathcal P(P_{Y|X},Q_Y)} I(X;Y),
\]
which matches the Shannon case. For \(s\in(0,1]\cup\{\infty\}\), it becomes
\[
\min_{P_X\in\mathcal P(P_{Y|X},Q_Y)}
\sum_x P_X(x)\,D_{1+s}\!\left(P_{Y|X}(\cdot|x)\|Q_Y\right),
\]
and is in general larger than the mutual-information benchmark. Once the rate exceeds the relevant resolvability, the optimal Rényi divergence vanishes at least exponentially fast [1707.00810].

The same resolvability machinery yields a complete characterization of the tradeoff between secret and non-secret message rates for the wiretap channel when leakage is measured by unnormalized Rényi divergence. In this sense, order-\(\alpha\) channel Rényi information is not only a generalized coding functional; it is also a secrecy criterion whose operational meaning changes sharply at \(\alpha=1\) [1707.00810].

## 4. Quantum channel Rényi information

Quantum generalizations begin with the choice of divergence. For the sandwiched Rényi divergence, a foundational theorem states that for
\[
\frac12<\alpha<\infty,
\]
and every completely positive, trace-preserving map \(\mathcal E\),
\[
D_\alpha(p\|o)\ge D_\alpha(\mathcal E(p)\|\mathcal E(o)).
\]
Equivalently, the sandwiched Rényi divergence obeys data processing for all \(\alpha\ge \tfrac12\). This monotonicity is the structural prerequisite for defining well-behaved quantum channel Rényi-information quantities in coding and strong-converse settings [1306.5358].

For a classical-quantum channel \(\mathscr W=\{W_x\}\subset\mathcal S(\mathcal H)\) and prior \(P\), noncommutative Rényi and Augustin informations are defined by
\[
I_\alpha^{r,(t)}(P,\mathscr W)
=
\inf_{\sigma\in\mathcal S(\mathcal H)}
D_\alpha^{(t)}(P\circ\mathscr W\|P\otimes \sigma),
\]
\[
I_\alpha^{a,(t)}(P,\mathscr W)
=
\inf_{\sigma\in\mathcal S(\mathcal H)}
\sum_{\omega\in\mathscr W}P(\omega)\,D_\alpha^{(t)}(\omega\|\sigma),
\]
where \((t)\in\{\,,*,\flat\}\) denotes the Petz, sandwiched, or log-Euclidean version. At \(\alpha=1\), both reduce to the Holevo information. The corresponding capacities satisfy
\[
C_{\alpha,\mathscr W}^{(t)}
=
\sup_P I_\alpha^{r,(t)}(P,\mathscr W)
=
\sup_P I_\alpha^{a,(t)}(P,\mathscr W)
\]
on the order ranges treated in the paper: Petz for \(\alpha\in[0,2]\), sandwiched for \(\alpha\in[1/2,\infty]\), and log-Euclidean for \(\alpha\in[0,\infty]\). Uniform equicontinuity, joint continuity in order and prior, and concavity of the scaled auxiliary functions \(E_0^{r,(t)}\) and \(E_0^{a,(t)}\) on \(s\in(-1,0)\) are then used to derive minimax formulas for strong-converse exponents, showing that the strong converse exponent can be attained by the best constant-composition code [1811.04218].

A distinct quantum line of work derives a Rényi-Holevo inequality from \(\alpha\)-\(z\)-Rényi relative entropies. For a measurement-induced classical channel \(W_{sq}^M\), it gives a one-letter upper bound on the classical Rényi divergence between the joint distribution and the product distribution, and hence on Sibson’s \(\alpha\)-mutual information:
\[
I_\alpha(p,W_{sq}^M)\le \frac{1}{\alpha-1}\log f_{sq}(p,\alpha).
\]
This bounds the order-\(\alpha\) channel capacity
\[
C_\alpha(W)=\max_p I_\alpha(p,W),
\]
and, through the Gallager function, also yields an upper bound on reliability functions for memoryless semi-quantum channels [2309.04539].

A further development replaces \(D_{\max}\) by the geometric Rényi divergence \(\widehat D_\alpha\), also called maximal Rényi divergence. Its channel extension satisfies a chain rule and hence an amortization collapse,
\[
\widehat D_\alpha^A(\mathcal N\|\mathcal M)=\widehat D_\alpha(\mathcal N\|\mathcal M),
\]
for \(\alpha\in(1,2]\). This produces single-letter upper bounds for channel discrimination exponents and for several channel capacities, sharpening the previously best-known \(D_{\max}\)-based bounds while remaining efficiently computable [1909.05758].

## 5. Multiple definitions, chain rules, and information combining

A recurring feature of the subject is that there is no single universally accepted conditional Rényi entropy. In the classical literature, Arimoto, Hayashi, Jizba–Arimitsu, and Cachin definitions are all used, while the quantum literature distinguishes several “up-arrow” and “down-arrow” versions built from sandwiched Rényi divergence. Likewise, the Shannon equalities
\[
I(A:B)=H(A)-H(A|B)=H(A)+H(B)-H(AB)
\]
do not generalize verbatim when \(\alpha\neq1\) [2004.14408][1912.06277].

Instead, one obtains order-dependent inequalities. For quantum sandwiched Rényi mutual information, if
\[
\frac{\alpha}{\alpha-1}=\frac{\beta}{\beta-1}+\frac{\gamma}{\gamma-1},
\]
then the paper proves decomposition rules such as
\[
I^\uparrow_\gamma(A;B)_\rho\ge H_\beta(B)_\rho-H^\downarrow_\alpha(B|A)_\rho
\]
when \((\alpha-1)(\beta-1)(\gamma-1)>0\), with the inequality reversed when the sign is negative. In the limit \(\alpha,\beta,\gamma\to1\), these collapse to the Shannon identity. The same work uses these inequalities to derive information-exclusion relations [1912.06277].

For channel combining and polar coding, the lack of a conventional chain rule is especially important. One paper supplies the missing link by proving chain rules that connect Hayashi and Arimoto conditional Rényi entropies via a power-reweighted distribution, allowing bounds for the “check-node” quantity \(H_\alpha^*(X_1+X_2\mid Y_1Y_2)\) to be translated into bounds for the complementary “variable-node” quantity \(H_\alpha^*(X_2\mid X_1+X_2,Y_1Y_2)\). In the special case \(\alpha=2\), it also gives the first optimal information-combining bounds with quantum side information [2305.02589].

A related classical treatment studies four conditional Rényi entropies and derives optimal BSC and BEC extremal bounds for
\[
H_\alpha^*(X_1+X_2\mid Y_1Y_2),
\]
provided an auxiliary combining map is convex or concave in each argument. For the Hayashi entropy, the convexity–concavity classification is complete: the relevant combining function is convex for \(0<\alpha<1\) and \(2<\alpha\le3\), concave for \(1<\alpha\le2\) and \(\alpha\ge3\), and linear at \(\alpha=2\) and \(\alpha=3\), where the BSC and BEC bounds coincide [2004.14408].

## 6. Polarization, leakage bounds, and scope of the concept

Rényi channel information is closely tied to polarization. Using the Jizba–Arimitsu/Golshani conditional entropy
\[
H_{\alpha}(X|Y)=\frac{1}{1-\alpha}\log \frac{\sum_{x,y} P_{X,Y}(x,y)^\alpha}{\sum_y P_Y(y)^\alpha},
\]
which satisfies a chain rule, one paper proves that for any binary-input DMC or binary source with side information and any \(\alpha\ge0\), the synthetic entropies
\[
H_N(i)=H_{\alpha}\!\left(U^i \mid Y^{1:N}, U^{1:i-1}\right)
\]
polarize to \(0\) or \(1\). The fraction of high-entropy synthetic channels converges to \(H_\alpha(X|Y)\), and the fraction of low-entropy synthetic channels converges to \(1-H_\alpha(X|Y)\). The same paper also shows that different Rényi orders can classify the same synthetic sub-channel in opposite ways: one order may see it as effectively uniform, another as effectively deterministic [1907.06423].

In side-channel analysis and guessing problems, order-\(\alpha\) conditional Rényi entropy also controls guessing performance. For a leakage channel \(X\to Y\), the paper on guessing advantage derives optimal lower bounds on the guessing moment \(G_\rho(X|Y)\) as a function of the Rényi-Arimoto conditional entropy \(H_\alpha(X|Y)\). In the small-leakage regime it proves that
\[
\Delta G_\rho(X;Y)\lesssim
\sqrt{
\frac{2\bigl(G_{2\rho}(M)-G_\rho(M)^2\bigr)}{\alpha}
}
\,\sqrt{\frac{\Delta H_\alpha(X;Y)}{\log e}},
\]
providing a non-asymptotic link between information leakage and guessing advantage [2401.17057].

A common misconception is that “channel Rényi information” denotes a single canonical replacement for Shannon mutual information. The literature instead defines several nonequivalent quantities—Sibson mutual information, Rényi capacity, minimax redundancy, Augustin information, Petz and sandwiched mutual informations, \(\alpha\)-\(z\) Holevo-type quantities, and geometric Rényi channel divergences—whose suitability depends on the operational task. This suggests that the unifying object is not one formula but a task-dependent Rényi-information framework, with data processing, minimax structure, and error-exponent semantics determining which version is appropriate in a given problem.

Source: https://www.emergentmind.com/topics/channel-renyi-information