---
title: Entropy Numbers in Banach Spaces
url: https://www.emergentmind.com/topics/entropy-numbers
type: topic
---

# Entropy Numbers in Banach Spaces

Entropy numbers are covering-based invariants of bounded sets and operators that quantify compactness in Banach and quasi-Banach geometry. In the operator form, they measure the smallest radius of balls needed to cover the image of the unit ball by a prescribed number of centers; in the set form, they encode the metric entropy of compact classes. Across the literature they serve as a common language for finite-dimensional embeddings, diagonal and multiplier operators, function classes, spectral estimates, and discretization theory. Their decay governs compactness, and in many settings it can be described sharply by power laws, logarithmic corrections, or exponential regimes [1802.00572], [2508.06542].

## 1. Definition and foundational properties

Let \(X\) and \(Y\) be Banach spaces, \(B_X\) the closed unit ball of \(X\), and let
\[
N\bigl(T(B_X),\;\varepsilon\,B_Y\bigr)
\]
be the smallest number of translates of \(\varepsilon B_Y\) needed to cover \(T(B_X)\). A standard dyadic definition of the \(n\)-th entropy number of a compact operator \(T:X\to Y\) is
\[
e_n(T):=\inf\bigl\{\varepsilon>0:\,N(T(B_X),\,\varepsilon B_Y)\le 2^{\,n-1}\bigr\}.
\]
For a bounded subset \(K\subset X\), one analogously sets
\[
e_n(K,X)=\inf\{\varepsilon>0:\,N(K,X,\varepsilon)\le 2^{\,n-1}\}.
\]
This formulation appears throughout the modern theory and is equivalent to asking for a covering by \(2^{n-1}\) balls of radius \(\varepsilon\) [1002.1377], [1802.00572].

Several structural facts recur across Banach and quasi-Banach settings. Entropy numbers are monotone,
\[
e_1(T)\ge e_2(T)\ge \cdots \ge 0,
\]
subadditive in the form
\[
e_{m+n-1}(S+T)\le e_m(S)+e_n(T)
\]
or, in a \(\varrho\)-normed target, \(e_{m+n-1}(S+T)^\varrho\le e_m(S)^\varrho+e_n(T)^\varrho\), and submultiplicative,
\[
e_{m+n-1}(RT)\le e_m(R)e_n(T).
\]
They also satisfy the compactness criterion
\[
T\text{ is compact}\quad\Longleftrightarrow\quad \lim_{n\to\infty} e_n(T)=0.
\]
In quasi-Banach spaces, entropy numbers are compared to approximation and Kolmogorov numbers by
\[
d_n(T)\le a_n(T),\qquad
e_n(T)\le c_p\Bigl(\frac1n\sum_{k=1}^n a_k(T)^p\Bigr)^{1/p},
\qquad
e_n(T)\le c_p\Bigl(\frac1n\sum_{k=1}^n d_k(T)^p\Bigr)^{1/p},
\]
and Carl’s inequality yields
\[
\Bigl(\prod_{m=1}^k|\lambda_m|\Bigr)^{1/k}\le \inf_{n\ge1}2^{\,n/(2k)}e_n(T),
\qquad |\lambda_k|\le \sqrt2\,e_k(T),
\]
linking covering behavior to eigenvalue decay [2508.06542].

When the target is a \(\gamma\)-Banach space, \(0<\gamma\le 1\), the first entropy number no longer coincides exactly with the operator norm. Instead one has the sharp estimate
\[
2^{1-1/\gamma}\|T\|\le e_1(T)\le \|T\|,
\]
and the constant \(2^{1-1/\gamma}\) is best possible. There are even examples with \(e_k(T)=2^{1-1/\gamma}\|T\|\neq 0\) for every \(k\) [1707.09200].

## 2. Canonical finite-dimensional models

The basic model problem is the identity embedding
\[
id:\ell_p^n\to \ell_q^n,\qquad 0<p,q\le \infty.
\]
Its entropy numbers exhibit the standard regime structure used throughout the subject. For \(0<p<q\le\infty\),
\[
e_k(id)\sim
\begin{cases}
1, & 1\le k\le \lfloor \log_2 n\rfloor,\\[4pt]
\bigl[\log_2(1+n/k)/k\bigr]^{1/p-1/q}, & \lfloor \log_2 n\rfloor<k\le n,\\[4pt]
2^{-(k-n)/n}\,n^{1/q-1/p}, & k>n,
\end{cases}
\]
whereas for \(0<q<p\le\infty\),
\[
e_k(id)\sim 2^{-(k-1)/n}\,n^{1/q-1/p}.
\]
Thus \(p<q\) yields a constant regime, then a polynomial regime, and finally an exponential tail; \(q<p\) yields a single exponential regime in \(k\). The proofs combine volume comparison, combinatorial packing by Hamming-type or constant-weight codes, interpolation, and in the quasi-Banach range Maurey’s empirical method [1802.00572].

This picture extends to finer scales. For finite-dimensional Lorentz spaces \(\ell_{p,u}^n\to \ell_{q,v}^n\), the asymptotics are sharp in all regimes. When \(0<p\neq q<\infty\), they coincide with the classical \(\ell_p^n\to \ell_q^n\) behavior, but extra logarithmic factors appear when one or both of \(p,q\) are infinite or when \(p=q\) and the secondary indices differ. In the intermediate range \(\log n\le k\le n\), the correct scale is expressed through
\[
\ell(k,n):=\frac{k}{\log(n/k+1)},
\]
and a key reduction identifies \(e_k(id:X\to Y)\) with the worst-case best \(s\)-term approximation error for \(s\approx k/\log(n/k)\) [2404.06058].

The noncommutative analogue is the Schatten embedding
\[
id:S_p^N\to S_q^N.
\]
Here the ambient dimension is \(N^2\), and the entropy profile is correspondingly different. For \(1\le n\le N^2\),
\[
e_n\bigl(id:S_p^N\to S_q^N\bigr)\asymp
\begin{cases}
1, & 1\le n\le N,\\[4pt]
(N/n)^{1/p-1/q}, & N\le n\le N^2.
\end{cases}
\]
For \(n\gg N^2\) one sees an exponential tail of the form \(2^{-2n/N^2}N^{1/q-1/p}\). In contrast to the classical \(\ell_p^N\to \ell_q^N\) setting, the matrix case has no logarithmic middle regime on \(1\le n\le N^2\) [1612.08105].

## 3. Critical operators and adaptive finite-rank approximation

A central development in the theory concerns “critical” operators, where standard truncation arguments leave a logarithmic gap. Two model classes are the summation operators on the infinite full binary tree and Volterra-type integral operators. For the tree model, with weight
\[
w(t)=(1+|t|)^{-B},\qquad B>1,
\]
one considers
\[
V:\ell_2(T,w)\to \ell_\infty(T),\qquad (Vf)(t)=\sum_{u\preceq t} w(u)f(u),
\]
and its adjoint
\[
V^*:\ell_1(T)\to \ell_2(T,w),\qquad (V^*p)(t)=\sum_{u\preceq t}p(u)=p(O(t)).
\]
For the Volterra model, with
\[
K^{(B)}(t,s)=(t-s)_+^{\,B-1/2}\,|\ln(t-s)_+|^{-B},\qquad 0\le s\le t\le r,
\]
one studies
\[
V_B:L_2(0,r)\to C[0,r],\qquad (V_Bf)(t)=\int_0^t K^{(B)}(t,s)f(s)\,ds,
\]
together with its adjoint \(V_B^*:M[0,r]\to L_2(0,r)\). The cases \(B=2\) for tree summation and \(B=1\) for the Volterra kernel are critical in the precise sense that previously known methods produced an extra logarithmic factor [1002.1377].

For the tree operator, the regular-case asymptotics are
\[
C_1\,n^{-(B-1/2)}\le e_n(V^*)\le C_2\,n^{-(B-1/2)},\qquad 1<B<2,
\]
\[
C_1\,n^{-1/2}\le e_n(V^*)\le C_2\,n^{-1/2}\ln n,\qquad B=2,
\]
\[
C_1\,n^{-1/2}(\ln n)^{1-B}\le e_n(V^*)\le C_2\,n^{-1/2}(\ln n)^{1-B},\qquad B>2.
\]
The critical theorem improves the upper bound at \(B=2\) to
\[
\max\{e_n(V),e_n(V^*)\}\le C\,n^{-1/2}.
\]
For the Volterra operator with
\[
K^{(1)}(t,s)=(t-s)^{-1/2}|\ln(t-s)|^{-1},
\]
one likewise has
\[
\max\{e_n(V_1),e_n(V_1^*)\}\le C\,n^{-1/2}.
\]
These results close the logarithmic gap left by the classical truncation procedure [1002.1377].

The method is an adaptive family approximation rather than a single fixed truncation. If \(\{T_\alpha\}_{\alpha\in\mathcal A}\) is a family of operators, then
\[
e_{\,n+\lceil \log_2|\mathcal A|\rceil+1}(T)
\le
\sup_{\alpha\in\mathcal A} e_n(T_\alpha)
+
\sup_{\|x\|\le 1}\inf_{\alpha\in\mathcal A}\|Tx-T_\alpha x\|.
\]
For tree summation, each \(p\in\ell_1(T)\) determines an \(n\)-essential subtree \(Y(p)\) by a stopping rule based on accumulated variation. Every such subtree satisfies
\[
\sum_{t\in Y}(1+|t|)^{-1}<n,\qquad \#Y\le 2(n+1),
\]
and the number of possible essential trees is at most \((4e)^n\). Truncating to \(Y\) yields a finite-rank operator \(A_Y\) with
\[
\|V^*-A_{Y(p)}\|_{\ell_1\to\ell_2}\le n^{-1/2},\qquad \operatorname{rank}A_Y\le 2(n+1).
\]
For Volterra operators, essential dyadic partitions of \([0,r]\) play the same role: they have at most \(2(n+1)\) intervals, at most \((4e)^n\) possibilities, and support finite-rank kernel approximations with error \(Cn^{-1/2}\) [1002.1377].

The failure of classical truncation is explicit. Truncation at level \(L\) gives rank \(\sim 2^L\) and error \(\sim L^{-1/2}\); balancing \(2^L\approx n\) yields
\[
e_n\approx n^{-1/2}(\ln n)^{1/2},
\]
so a logarithmic gap remains. Adaptive truncation removes precisely this \(\sqrt{\ln n}\) loss by tailoring the finite-rank approximation to each input [1002.1377].

## 4. Diagonal, multiplier, and adjoint phenomena

Diagonal operators on sequence spaces provide another major testing ground. Given a nonincreasing sequence \(\sigma=(\sigma_k)_{k\ge 1}\) with \(\sigma_k\to 0\), define
\[
D_\sigma:\ell_p\to \ell_q,\qquad D_\sigma(x_1,x_2,\dots)=(\sigma_1x_1,\sigma_2x_2,\dots).
\]
For \(0<p<q\le \infty\), if \(1/s=1/p-1/q\), then one has an upper product bound
\[
\varepsilon_n(D_\sigma)\lesssim
\sup_{k\ge 1}
k^{-1/s}
\Bigl[(\sigma_1+k^{1/s}\sigma_1)\cdots(\sigma_k+k^{1/s}\sigma_k)\Bigr]^{1/k}
\,n^{-1/k}.
\]
Under the exponential decay condition
\[
\exists b>1:\sup_{1\le k\le n}\frac{\sigma_n b^n}{\sigma_k b^k}<\infty,
\]
this becomes sharp, and
\[
\varepsilon_n(D_\sigma)\asymp
\sup_{k\ge1}k^{-1/s}\,(\sigma_1\cdots \sigma_k/n)^{1/k}.
\]
For \(0<q<p\le\infty\), with \(1/q=1/p+1/r\) and tail sequence
\[
\tau_k:=\Bigl(\sum_{n=k}^\infty \sigma_n^r\Bigr)^{1/r},
\]
two regimes appear. Under the “at least polynomial” hypothesis \((ALP)\), one gets
\[
\varepsilon_n(D_\sigma)\asymp
\sup_{k\ge1}k^{1/r}(\sigma_1\cdots \sigma_k/n)^{1/k},
\]
whereas under the “at most polynomial” hypothesis \((AMP)\),
\[
\varepsilon_n(D_\sigma)\asymp \tau_{\lfloor \log_2 n\rfloor+1}.
\]
The proof mechanism is decomposition into a finite-dimensional truncation plus tail, followed by volume estimates in dimension \(k\) and optimization in \(k\) [1903.00541].

A complementary line of work concerns approximation by finite-dimensional compressions. If
\[
T_n=Q_nTP_n,\qquad \|P_n\|\le 1,\quad \|Q_n\|\le 1,
\]
and \(T_nx\to Tx\) in norm for each \(x\), then for each fixed \(k\),
\[
\lim_{n\to\infty} e_k(T_n)=e_k(T)
\]
when the target is reflexive, and also in certain dual settings. In separable Hilbert spaces this leads to the exact adjoint symmetry
\[
e_n(T)=e_n(T^*)=e_n(|T|),
\]
which gives a complete affirmative answer to Carl’s question in that setting [1703.01418].

For Fourier multiplier operators,
\[
T_\lambda f=\sum_{\mathbf m\in\mathbb N_0^d}\lambda_{\mathbf m}\langle f,\phi_{\mathbf m}\rangle \phi_{\mathbf m},
\]
defined with respect to a bounded orthonormal system \(\{\phi_{\mathbf m}\}\subset L^2(\Omega)\cap L^\infty(\Omega)\), the entropy problem is transferred to finite-dimensional diagonal blocks. If \(\lambda_{\mathbf m}=\psi(|\mathbf m|)\) with \(\psi\) nonincreasing and \(n=\dim T_N\), then
\[
e_k(T_\lambda:L^p\to L^q)\ge C\,2^{-k/n}|\psi(N)|\,V_n,
\]
while the upper bounds are controlled by a factor \(\Gamma_{p,q}(n)\). In concrete Sobolev-type cases,
\[
e_k(T_\lambda)\asymp k^{-s/d},
\]
and if \(\lambda_{\mathbf m}=|\mathbf m|^{-y}(\log_2|\mathbf m|)^{-s}\) with \(y>d/2\), then
\[
e_n(T_\lambda:L^p\to L^q)\asymp n^{-y/d}(\log n)^{-s}.
\]
For Gevrey-type multipliers \(\lambda_{\mathbf m}=\exp(-\vartheta |\mathbf m|^r)\), \(0<r<1\),
\[
e_n(T_\lambda:L^2\to L^2)\asymp \exp\bigl(-C\,n^{r/(d+r)}\bigr).
\]
These rates are order-sharp in the main exponents of \(n\) and \(\log n\) [2107.04093].

## 5. Function classes and metric entropy

Entropy numbers of function classes often reveal an intrinsic dimension that is not visible from the ambient domain. For ridge functions
\[
R_p^\alpha=\{f:B^d\to \mathbb R:\ f(x)=g(a\cdot x),\ g\in \operatorname{Lip}^\alpha([-1,1]),\ \|a\|_p\le 1\},
\]
measured in \(L_\infty(B^d)\), the univariate reference class
\[
B^\alpha=\{g:[-1,1]\to\mathbb R:\ |g|_{\operatorname{Lip}^\alpha}\le 1\}
\]
satisfies
\[
e_k(B^\alpha,L_\infty[-1,1])\asymp C_\alpha\,k^{-\alpha}.
\]
For \(R_p^\alpha\), one has the regime picture: \(e_n(R_p^\alpha,L_\infty(B^d))\asymp 1\) for \(n\lesssim d\); an intermediate decay controlled by the direction-net term up to a logarithmic-in-\(d\) threshold; and
\[
e_n(R_p^\alpha,L_\infty(B^d))\asymp C(d,p,\alpha)\,n^{-\alpha}
\]
once \(n\gtrsim C\,d\log d\). By contrast,
\[
e_n(\operatorname{Lip}^\alpha(B^d),L_\infty)\asymp n^{-\alpha/d},
\]
so the ridge class has one-dimensional entropy asymptotics for large \(n\), despite living on a high-dimensional domain [1311.2005].

Mixed-smoothness classes on the torus display a different but equally structured behavior. For
\[
W_q^{a,b}
=
\Bigl\{f:\|\Delta_s f\|_{L_q}\le 2^{-a|s|_1}(|s|_1)^{-(d-1)b}\Bigr\},
\]
sharp estimates include
\[
e_k(W_1^{a,b},L_p)\asymp k^{-a}(\log k)^{a+b},\qquad 1<p<\infty,\ d=2,
\]
\[
e_k(W_1^{a,b},L_\infty)\asymp k^{-a}(\log k)^{a+b+1/2},
\]
\[
e_k(W_\infty^{a,b},L_p)\asymp k^{-a}(\log k)^{a+b-1/2},\qquad d=2,
\]
and for general \(d\ge 2\),
\[
e_k(W_q^{a,b},L_q)\asymp k^{-a}(\log k)^{(d-1)(a+b)},\qquad 1<q<\infty.
\]
Here the upper bounds are obtained by a two-step strategy: first derive best \(m\)-term approximation estimates with respect to a suitable dictionary, then convert them into entropy estimates through a general inequality of the form
\[
\sigma_m(F,\mathcal D)_X\le C m^{-r}\ \Longrightarrow\
e_k(F,X)\le C'(r)\,k^{-r}\Bigl(\log\frac{2N}{k}\Bigr)^r,\qquad k\le N.
\]
This is one of the clearest examples of nonlinear approximation feeding directly into entropy asymptotics [1602.08712].

For finite-dimensional subspaces \(X_N\subset L_\infty(\Omega)\), the entropy of the \(L_p\)-unit ball
\[
X_N^p:=\{f\in X_N:\|f\|_{L_p(\Omega)}\le 1\}
\]
controls sampling discretization. Under the Nikol’skii-type assumptions
\[
\|f\|_\infty\le (K_1N)^{1/2}\|f\|_{L_2(\Omega)},
\qquad
\|f\|_\infty\le K_2\|f\|_{L_p(\Omega)}\log N,
\]
one has
\[
e_k(X_N^p,L_\infty(\Omega))
\le
C_p\,(K_1K_2\log N)^{1/p-1/2}
\begin{cases}
(N/k)^{1/p}, & 1\le k\le N,\\[4pt]
2^{-k/N}, & k>N.
\end{cases}
\]
Combined with a conditional discretization theorem, this yields equal-weight Marcinkiewicz discretization with
\[
m\le C_p(a,\varepsilon)\,K_1N(\log N)^3
\]
under the growth condition \(\log K_1<a\log N\), and a weighted version for arbitrary \(X_N\subset L_p(\Omega)\) with
\[
m\le C_p(\varepsilon)\,N(\log N)^3.
\]
The Rademacher example shows that the endpoint condition involving \(\log N\) is genuinely needed for the sharp first-block rate when \(1<p<2\) [2001.10636].

Entropy numbers also interact with Minkowski dimension. For a connected totally bounded set \(K\subset X\),
\[
\dim_B K<\infty
\quad\Longleftrightarrow\quad
\limsup_{n\to\infty}[e_n(K)]^{1/n}<1.
\]
If \(P:X\to Y\) is an \(m\)-homogeneous polynomial, then
\[
\dim_B(P(B_X))<\infty
\quad\Longleftrightarrow\quad
\operatorname{span}P(B_X)\text{ is finite-dimensional}.
\]
This fails for holomorphic maps: there are holomorphic \(f\) with \(f(B_X)\) of finite box dimension but infinite-dimensional linear span. Moreover, if \(f\) is holomorphic near \(x\) and \(P_m(x)\) are the homogeneous Taylor coefficients, finite box dimension of \(f(B_X(x,\varepsilon))\) implies
\[
\bigl(e_n(P_m(x)(B_X(0,\varepsilon)))\bigr)_{n=1}^\infty\in \ell^p
\qquad\text{for every }p>1.
\]
The relation is mediated by explicit inequalities comparing the entropy of \(f\) with that of its Taylor polynomials [2401.12059].

## 6. Methods, recurring themes, and significance

A small number of proof mechanisms recur across the subject. Volume comparison yields lower bounds by comparing \(\operatorname{vol}(T(B_X))\) with \(\operatorname{vol}(\varepsilon B_Y)\); this is decisive for \(\ell_p^n\)-embeddings, diagonal operators, Lorentz embeddings, and Schatten classes. Combinatorial constructions—Hamming codes, constant-weight codes, sparse support packings, and separated sets—produce the sharp intermediate regimes. Interpolation and duality pass between endpoint cases and general parameters. Nonlinear approximation provides another route: greedy and \(m\)-term methods convert approximation rates into entropy rates, and Temlyakov’s refinement of Talagrand’s theorem replaces \(\log n\) by \(\log(2n/k)\) in entropy bounds for octahedra in uniformly smooth spaces [1802.00572], [2008.13030].

Two distinctions are especially important. First, the entropy profile need not reflect ambient dimension in a naive way. Ridge-function classes can have asymptotic rate \(n^{-\alpha}\) independently of \(d\), while full multivariate Lipschitz classes have \(n^{-\alpha/d}\) decay [1311.2005]. Second, fixed truncation is not always structurally adequate. In critical tree and Volterra problems, the optimal \(n^{-1/2}\) upper bound becomes visible only after replacing one truncation by a family of input-dependent finite-rank approximants indexed by essential trees or partitions [1002.1377].

These results feed directly into neighboring theories. Entropy numbers provide upper and lower bounds for Kolmogorov widths and Gelfand widths, control covering numbers used in Johnson–Lindenstrauss lower bounds, enter statistical learning through uniform convergence estimates for hypothesis classes, and appear in compressed sensing via the same net arguments used for sparse recovery [1802.00572]. Through Carl’s and Weyl’s inequalities they connect covering geometry to spectral theory: small entropy numbers force rapid eigenvalue decay, while good finite-rank approximation controls products and \(\ell^p\)-sums of eigenvalues [2508.06542]. In the noncommutative setting they yield lower bounds for low-rank matrix recovery through their relation to Gelfand numbers [1612.08105].

Taken together, the modern theory presents entropy numbers as a unifying invariant across geometric functional analysis, approximation theory, operator theory, and information-based complexity. Their most characteristic features are the coexistence of universal formal properties with highly problem-specific asymptotics, and the fact that apparently small structural changes—critical exponents, infinite secondary indices, quasi-Banach targets, or nonlinear model classes—can change the entropy scale from a pure power law to a logarithmically corrected or exponential regime.

Source: https://www.emergentmind.com/topics/entropy-numbers