---
title: Kolmogorov Numbers in Operator Approximation
url: https://www.emergentmind.com/topics/kolmogorov-numbers
type: topic
---

# Kolmogorov Numbers in Operator Approximation

Searching arXiv for recent and foundational papers on Kolmogorov numbers to ground the article in published work.
Kolmogorov numbers are an \(s\)-number sequence associated with a bounded linear operator \(T:X\to Y\). They quantify the smallest worst-case error obtained by approximating the image of the unit ball \(T(B_X)\) by subspaces of \(Y\) of dimension \(<n\). For compact operators they form a nonincreasing sequence tending to zero; in the Hilbert-space case they coincide with the singular values; and, in a communication-theoretic formulation, they are exactly the jump points of the degrees-of-freedom function \(N(\epsilon)\) of a linear channel [1103.3510].

## 1. Definition and geometric content

Let \(X\) and \(Y\) be Banach spaces and let \(T\colon X\to Y\) be bounded. If \(Q_S^Y\colon Y\to Y/S\) denotes the quotient map onto the quotient by a closed subspace \(S\subset Y\), then the \(n\)th Kolmogorov number is
\[
d_n(T)
=
\inf_{\dim S<n}\,\|Q_S^Y\,T\|
=
\inf_{\dim S<n}\,\sup_{\|x\|_X\le1}\,\inf_{y\in S}\|Tx-y\|_Y.
\]
Equivalently, \(d_n(T)\) is the worst-case distance of \(T(B_X)\) to an \((n-1)\)-dimensional subspace of \(Y\). This makes explicit that Kolmogorov numbers are codomain-side approximation quantities: one fixes a low-dimensional subspace of \(Y\) and measures how well the whole image \(T(B_X)\) can be represented inside it [1103.3510].

A closely related finite-dimensional approximation formula is
\[
d_{n+1}(T)
=
\inf_{\psi_1,\dots,\psi_n\in Y}
\sup_{\|x\|_X\le1}
\inf_{a_1,\dots,a_n\in\mathbb R}
\Big\|Tx-\sum_{i=1}^n a_i\psi_i\Big\|_Y.
\]
This version makes clear that the problem can be read as selecting \(n\) trial vectors in the target space and minimizing the maximal residual over the unit ball of the domain [1103.3510].

The same basic definition extends to quasi-Banach spaces:
\[
d_n(T)
=
\inf_{\dim U<n}\,\sup_{\|x\|_X\le1}\,\inf_{y\in U}\|Tx-y\|_Y.
\]
In that setting, too, the quantity measures approximation of \(T(B_X)\) by finite-dimensional subspaces of the target [2508.06542].

## 2. Position within the theory of \(s\)-numbers

Kolmogorov numbers are one of the classical \(s\)-numbers in the sense of Pietsch. In the Banach-space setting they satisfy Pietsch’s axioms SN1–SN5 and therefore define a valid \(s\)-number sequence. In particular,
\[
\|T\|=d_1(T)\ge d_2(T)\ge\cdots\ge0,
\]
they obey the ideal property
\[
d_n(BTA)\le \|B\|\,d_n(T)\,\|A\|,
\]
they are normalized by \(d_n(I)=1\) on any Banach space, and they vanish once the rank drops below the index:
\[
\operatorname{rank}T<n\quad\Longrightarrow\quad d_n(T)=0.
\]
These properties place them in the same structural framework as approximation, Gelfand, Bernstein, and Hilbert numbers [1103.3510; 2405.05509].

In quasi-Banach spaces, the structural picture persists with the expected modifications. One has monotonicity,
\[
\|T\|=d_1(T)\ge d_2(T)\ge\cdots\ge0,
\]
quasi-additivity,
\[
d_{m+n-1}(S+T)\le C_Y\bigl(d_m(S)+d_n(T)\bigr),
\]
multiplicativity,
\[
d_{m+n-1}(RS)\le d_m(R)\,d_n(S),
\]
the rank property, normalization for identities, and the ideal property
\[
d_n(RTU)\le \|R\|\,d_n(T)\,\|U\|.
\]
A central characterization is
\[
T\text{ is compact }\iff \lim_{n\to\infty} d_n(T)=0,
\]
which makes Kolmogorov numbers a compactness scale as well as an approximation scale [2508.06542].

A recent comparison theorem gives a sharp product-type bound between small and large \(s\)-numbers:
\[
\max\{c_n(T),d_n(T)\}
\le
n\Bigl(\prod_{k=1}^n h_k(T)\Bigr)^{1/n}.
\]
Here \(c_n\) are Gelfand numbers and \(h_n\) are Hilbert numbers. This shows that Kolmogorov numbers cannot decay, on average, much faster than the Hilbert numbers up to the optimal factor \(n\) [2405.05509].

## 3. Hilbert-space specialization and relations to neighboring quantities

When \(X\) and \(Y\) are Hilbert spaces and \(T\) is compact, the \(s\)-number theory collapses to the singular-value theory. In that case,
\[
d_n(T)=\sigma_n(T),
\]
the \(n\)th singular value of \(T\). This is one of the main reasons Kolmogorov numbers are viewed as the Banach-space analogue of singular values [1103.3510].

Several neighboring quantities recur throughout the literature.

| Quantity | Notation | Relation stated in the sources |
|---|---|---|
| Approximation numbers | \(a_n(T)\) | \(d_n(T)\le a_n(T)\) |
| Gelfand numbers of the adjoint | \(c_n(T')\) | \(d_n(T)=c_n(T')\) |
| Hilbert numbers | \(h_n(T)\) | \(h_n(T)\le d_n(T)\) |
| Entropy numbers | \(e_n(T)\) | For compact Hilbert-to-Banach maps, decay of \(d_n\) and \(e_n\) is equivalent in several regimes |
| Interpolation widths | \(I_n(H,L_p(\mu))\) | \(d_n(T)\le a_n(T)\le I_n(H,L_p(\mu))\) |

These comparisons clarify a common source of confusion: Kolmogorov numbers are not, in general, the same as approximation numbers, entropy numbers, or interpolation widths, even though they are often comparable [1606.05500; 2405.05509].

For compact operators \(S:H\to F\) from a Hilbert space into a Banach space, Steinwart proves that for \(p\in(0,2)\),
\[
d_n(S)\prec n^{-1/p}
\quad\Longleftrightarrow\quad
e_n(S)\prec n^{-1/p},
\]
and, if \(F\) has the metric-extension property, analogous equivalences hold for \(p\in(2,\infty)\) as well [1606.05500]. In the setting of RKHS embeddings into \(L_\infty\), the same paper also exhibits a genuine half-power separation between Kolmogorov widths and interpolation widths: if \(\lambda_n(T_k)\asymp n^{-2/\alpha}\) with \(\alpha\in(0,2)\), then
\[
d_n(H\to L_\infty)\asymp n^{-1/\alpha},
\qquad
I_n(H,L_\infty)\succcurlyeq n^{-1/\alpha+1/2}.
\]
The same \(n^{1/2}\) gap is attained for multidimensional Sobolev embeddings \(W^m(X)\to L_\infty(X)\) [1606.05500].

A separate terminological issue appears in a recent Lie-group paper, where the covering number \(\mathcal C(\varepsilon,L)\) is described as an “entropy Kolmogorov number,” while the standard Kolmogorov numbers remain the quantities \(d_n(L)\). In that setting the paper states
\[
d_n(L)\le e_n(L)\le 2\,d_n(L),
\]
so the distinction between covering, entropy, and Kolmogorov numbers is preserved even when the terminology is broadened [2602.02305].

## 4. Communication channels and the degrees-of-freedom function

A linear communication channel may be modeled by a compact linear operator \(T:X\to Y\), where \(\|x\|_X\le1\) is a transmitter power constraint and \(\epsilon>0\) is a receiver noise threshold below which signals cannot be distinguished. In that setting the number of degrees of freedom at level \(\epsilon\) is defined by
\[
N(\epsilon)
=
\min\Bigl\{
N\in\mathbb N_0:
\exists\,\psi_1,\dots,\psi_N\in Y
\text{ such that }
\sup_{\|x\|_X\le1}
\inf_{a_1,\dots,a_N}
\Big\|Tx-\sum_{i=1}^N a_i\psi_i\Big\|_Y
\le \epsilon
\Bigr\}.
\]
Operationally, \(N(\epsilon)\) is the number of linearly independent signals that may be communicated through the channel at noise level \(\epsilon\) [1103.3510].

The function \(\epsilon\mapsto N(\epsilon)\) has the basic properties expected of a degrees-of-freedom count. One has \(N(\epsilon)=0\) for all \(\epsilon\ge\|T\|\), the map is nonincreasing, and it has only finitely many jumps in any finite \(\epsilon\)-interval. Unless \(T=0\), \(N(\epsilon)\nearrow\infty\) as \(\epsilon\searrow0\) [1103.3510].

If \(\sigma_n\) denotes the \(n\)th jump point of \(N(\epsilon)\), characterized by
\[
\sup_{\epsilon>\sigma_n}N(\epsilon)=n-1,
\qquad
\inf_{\epsilon<\sigma_n}N(\epsilon)\ge n,
\]
then the central theorem is
\[
\sigma_n=d_n(T).
\]
Thus Kolmogorov numbers are exactly the discontinuity thresholds of the degrees-of-freedom function. This gives an operator-theoretic quantity an immediate communication-theoretic interpretation: the \(n\)th Kolmogorov number is the noise level at which the effective dimensionality of the channel drops from at least \(n\) to at most \(n-1\) [1103.3510].

## 5. Asymptotic theories for embeddings and finite-dimensional models

A major part of the literature studies asymptotics of \(d_n(T)\) for concrete embeddings. For weighted Sobolev-type embeddings of Besov and Triebel–Lizorkin spaces with polynomial weights,
\[
\operatorname{id}:
B^{s_1}_{p_1,q_1}(\mathbb R^d,w_\alpha)\to B^{s_2}_{p_2,q_2}(\mathbb R^d)
\]
and analogous \(F\)-space variants, the non-limiting case is governed by
\[
\delta=s_1-s_2-d\Bigl(\frac1{p_1}-\frac1{p_2}\Bigr)>0,
\qquad
\mu=\min(\alpha,\delta).
\]
The sharp asymptotics are of the form
\[
d_n\sim n^{-\gamma}
\quad\text{or}\quad
d_n\sim n^{-\kappa},
\]
with six parameter regimes determined by \(p_1,p_2,\mu,d\), and the same rates remain valid for the corresponding Besov/Triebel–Lizorkin combinations under the stated quasi-normability restrictions [1102.0677; 1105.5499].

The proofs in these weighted function-space problems proceed by wavelet discretization, reduction to sequence-space embeddings, and use of sharp finite-dimensional \(\ell_p\to\ell_q\) estimates. One representative finite-dimensional identity is
\[
d_k\bigl(\operatorname{id}:\ell_p^n\to\ell_q^n\bigr)
=
(n-k+1)^{1/q-1/p}
\qquad
(q\le p\le\infty),
\]
while other parameter ranges exhibit piecewise asymptotics involving logarithmic factors, \(k^{-1/2}\)-decay, or interpolation between two regimes [2508.06542]. These finite-dimensional estimates are also used explicitly in the weighted Sobolev papers through dyadic decomposition and operator-ideal arguments [1105.5499].

For Schatten-class embeddings
\[
T=\operatorname{id}:\mathcal S_p^N\to\mathcal S_q^N,
\]
Prochno and Strzelecki obtain asymptotically sharp two-sided bounds, up to constants depending only on \(p,q\), across several parameter regions. For example, in the “upper-triangle” regime \(1\le q\le p\le\infty\),
\[
d_n(T)\asymp_{p,q}\max\Bigl\{1,\frac{N^2-n+1}{N}\Bigr\}^{1/q-1/p}.
\]
In the quasi-Banach range \(0<p\le q\le1\), the same paper states
\[
d_n\bigl(\mathcal S_p^N\hookrightarrow \mathcal S_q^N\bigr)=1,
\qquad
1\le n\le N^2.
\]
The analysis relies on duality with Gelfand numbers, volume/Dvoretzky arguments, interpolation, and explicit construction of quotients of \(\mathcal S_q^N\) of controlled norm [2103.13050].

Another strand concerns embeddings of Besov spaces of dominating mixed smoothness into \(L_\infty\). The abstract of "Kolmogorov Numbers of Embeddings of Besov Spaces of Dominating Mixed Smoothness into \(L_\infty\)" announces two-sided sharp estimates for
\[
S^t_{p,q}B((0,1)^d)\hookrightarrow L_\infty((0,1)^d),
\]
placing mixed-smoothness problems within the same general width-theoretic framework [1411.6926].

## 6. Numerical computation and further analytic settings

For computation, a practical scheme is available when the domain \(X\) has a Schauder basis \((\phi_i)_{i\ge1}\). If
\[
S_m=\operatorname{span}\{\phi_1,\dots,\phi_m\},
\qquad
T_m=T|_{S_m}:S_m\to Y,
\]
and \(\sigma_{n,m}=d_n(T_m)\), then for each fixed \(n\),
\[
\sigma_{n,m}\le \sigma_n
\quad\text{and}\quad
\lim_{m\to\infty}\sigma_{n,m}=\sigma_n
\]
once \(m\) is large enough that the \(n\)th Kolmogorov number of the truncation exists. This leads to a finite-dimensional approximation procedure: choose \(m\), assemble the matrix of \(T_m\), compute \(d_n(T_m)\), and increase \(m\) until the values stabilize. In the Hilbert-space case this reduces to extracting the \(n\)th singular value of the finite matrix, or equivalently diagonalizing the associated Gram matrix [1103.3510].

The same approximation-theoretic viewpoint now appears in newer geometric settings. For a compact Lie group \(G\), a left-invariant positive symmetric trace-class integral kernel \(K\) induces an RKHS \(\mathcal H_K\subset C(G)\), and one studies the embedding
\[
\operatorname{id}:\mathcal H_K\hookrightarrow C(G).
\]
In that context, the paper on compact Lie groups defines covering numbers \(\mathcal C(\varepsilon,\operatorname{id})\), entropy numbers \(e_n(\operatorname{id})\), and Kolmogorov numbers \(d_n(\operatorname{id})\), and proves two-sided asymptotic estimates for \(e_n(\operatorname{id})\asymp d_n(\operatorname{id})\) from spectral assumptions on the group Fourier symbol of the kernel operator. Under a trace-decay hypothesis \(\operatorname{Tr}[\sigma_T(\xi)]\le b_T d_\xi \langle\xi\rangle^{-\beta}\) with \(\beta>n=\dim G\), the paper derives power-law decay of \(e_m(\operatorname{id})\) up to logarithmic factors; under a determinant-decay hypothesis, it derives logarithmic asymptotics of the form \((\log m)^{-\gamma/n}\) in the respective regime [2602.02305].

Taken together, these results show that Kolmogorov numbers serve simultaneously as abstract \(s\)-numbers, as singular-value analogues on Banach and quasi-Banach spaces, as operational thresholds in communication channels, and as sharp asymptotic invariants for embeddings in function spaces, matrix ideals, RKHS theory, and harmonic analysis on compact groups [1103.3510; 1606.05500; 2103.13050; 2602.02305].

Source: https://www.emergentmind.com/topics/kolmogorov-numbers