---
title: Mean Information Dimension
url: https://www.emergentmind.com/topics/mean-information-dimension
type: topic
---

# Mean Information Dimension

Searching arXiv for recent and foundational papers on mean information dimension, rate-distortion dimension, and metric mean dimension.
Mean information dimension is the dynamical or process-level analogue of Rényi’s information dimension, defined by the small-distortion slope of a rate-distortion function or, equivalently in the stationary process setting, by the growth rate of the entropy rate of a uniformly quantized process. In the mean-dimension literature it is usually formalized as the **rate-distortion dimension** of a measure-preserving dynamical system, while in stochastic-process theory it appears as the **information dimension rate**; the two viewpoints are linked explicitly in the high-resolution regime [1901.05623], [1702.00645]. In symbolic dynamics and mean dimension theory, mean information dimension functions as an information-theoretic counterpart to metric mean dimension and mean Hausdorff dimension, and in several important classes of systems it admits exact entropy formulas [1910.00844].

## 1. Formal definitions and basic quantities

Let $(\mathcal{X},T)$ be a compact metrizable dynamical system, let $d$ be a compatible metric, and define the Bowen metrics
$$
d_N(x,y)=\max_{0\le n<N} d(T^n x,T^n y).
$$
The scale-$\varepsilon$ dynamical covering complexity is
$$
S(\mathcal{X},T,d,\varepsilon)=\lim_{N\to\infty}\frac{\log \#(\mathcal{X},d_N,\varepsilon)}{N},
$$
where $\#(\mathcal{X},d_N,\varepsilon)$ is the $\varepsilon$-covering number. The upper and lower metric mean dimensions are then
$$
\overline{\mathrm{mdim}}_{\mathrm{M}}(\mathcal{X},T,d)
=\limsup_{\varepsilon\to 0}\frac{S(\mathcal{X},T,d,\varepsilon)}{\log(1/\varepsilon)},
\qquad
\underline{\mathrm{mdim}}_{\mathrm{M}}(\mathcal{X},T,d)
=\liminf_{\varepsilon\to 0}\frac{S(\mathcal{X},T,d,\varepsilon)}{\log(1/\varepsilon)}.
$$
When the two coincide, the common value is denoted $\mathrm{mdim}_{\mathrm{M}}(\mathcal{X},T,d)$ [1901.05623].

For a $T$-invariant Borel probability measure $\mu$, the rate-distortion function at distortion level $\varepsilon>0$ is
$$
R(d,\mu,\varepsilon)
=
\inf \frac{I(X;Y)}{N},
$$
where $X\sim\mu$, $Y=(Y_0,\dots,Y_{N-1})$ takes values in $\mathcal{X}^N$, and
$$
\mathbb{E}\left(\frac{1}{N}\sum_{n=0}^{N-1} d(T^nX,Y_n)\right)<\varepsilon.
$$
The associated upper and lower rate-distortion dimensions are
$$
\overline{\mathrm{rdim}}(\mathcal{X},T,d,\mu)
=\limsup_{\varepsilon\to 0}\frac{R(d,\mu,\varepsilon)}{\log(1/\varepsilon)},
\qquad
\underline{\mathrm{rdim}}(\mathcal{X},T,d,\mu)
=\liminf_{\varepsilon\to 0}\frac{R(d,\mu,\varepsilon)}{\log(1/\varepsilon)}.
$$
When they coincide, the common value is written $\mathrm{rdim}(\mathcal{X},T,d,\mu)$ [1901.05623].

In the process-theoretic literature, the same small-distortion slope is called the **information dimension rate** or **mean information dimension**. For a stationary process $\{X_t\}$, it is given by
$$
d(\{X\})=\lim_{m\to\infty}\frac{H'(\{[X_t]_m\})}{\log m},
$$
when the limit exists, where $[X]_m=\lfloor mX\rfloor/m$ and $H'$ denotes entropy rate [1702.00645].

| Quantity | Definition type | Interpretation |
|---|---|---|
| $\mathrm{mdim}_{\mathrm{M}}$ | Covering growth under $d_N$ | Dimension per iterate |
| $\mathrm{rdim}$ | Small-distortion slope of $R(d,\mu,\varepsilon)$ | Mean information dimension |
| $\mathrm{MRID}$ | Shrinking-partition entropy growth | Dynamical Rényi information dimension |

In the cited works, logarithms are taken in base $2$ [1901.05623], [1910.00844].

## 2. Relation to mean dimension and geometric dimension theories

Mean information dimension is part of a triad consisting of metric mean dimension, mean Hausdorff dimension, and rate-distortion dimension. Mean dimension itself is a topological invariant defined through width dimension:
$$
\mathrm{mdim}(\mathcal{X},T)
=
\lim_{\varepsilon\to 0}\left(\lim_{N\to\infty}\frac{\mathrm{Widim}_\varepsilon(\mathcal{X},d_N)}{N}\right),
$$
and is independent of the compatible metric [1901.05623].

The basic comparison inequalities place rate-distortion dimension below metric mean dimension:
$$
\overline{\mathrm{rdim}}(\mathcal{X},T,d,\mu)\le
\overline{\mathrm{mdim}}_{\mathrm{M}}(\mathcal{X},T,d),
\qquad
\underline{\mathrm{rdim}}(\mathcal{X},T,d,\mu)\le
\underline{\mathrm{mdim}}_{\mathrm{M}}(\mathcal{X},T,d)
$$
for all invariant measures $\mu$ [1901.05623]. When the metric has **tame growth of covering numbers**, meaning that for every $\delta>0$,
$$
\lim_{\varepsilon\to 0}\varepsilon^\delta \log \#(\mathcal{X},d,\varepsilon)=0,
$$
one also has the reverse-type variational bound
$$
\mathrm{mdim}_{\mathrm{H}}(\mathcal{X},T,d)
\le
\sup_{\mu\in\mathscr{M}^T(\mathcal{X})}
\underline{\mathrm{rdim}}(\mathcal{X},T,d,\mu),
$$
where $\mathrm{mdim}_{\mathrm{H}}$ is mean Hausdorff dimension [1901.05623]. This places mean information dimension at the interface between information theory and geometric measure theory.

A central result is the **double variational principle**: if $(\mathcal{X},T)$ has the marker property, then
$$
\mathrm{mdim}(\mathcal{X},T)
=
\min_{d\in\mathscr{D}(\mathcal{X})}
\sup_{\mu\in\mathscr{M}^T(\mathcal{X})}
\overline{\mathrm{rdim}}(\mathcal{X},T,d,\mu)
=
\min_{d\in\mathscr{D}(\mathcal{X})}
\sup_{\mu\in\mathscr{M}^T(\mathcal{X})}
\underline{\mathrm{rdim}}(\mathcal{X},T,d,\mu),
$$
and the minimum over metrics is attained [1901.05623]. Thus mean dimension can be recovered as a minimax of mean information dimension over compatible metrics and invariant measures.

The marker property is the condition that for every $N>0$ there exists an open set $U$ such that
$$
\mathcal{X}=\bigcup_{n\in\mathbb{Z}}T^{-n}U,
\qquad
U\cap T^{-n}U=\varnothing\quad \forall\,1\le n\le N.
$$
It implies freeness and is satisfied by free minimal systems and their extensions [1901.05623].

## 3. Dynamical Rényi information dimension

A distinct but closely related formulation replaces mutual information by measure-theoretic entropy of small-diameter partitions. For a finite Borel partition $\mathcal{P}$, denote by $h_\mu(T,\mathcal{P})$ its Kolmogorov–Sinai entropy rate. The upper and lower **mean Rényi information dimensions** are
$$
\overline{\mathrm{MRID}}(X,T,\mu,d)
=
\limsup_{\varepsilon\to 0}
\frac{1}{\log(1/\varepsilon)}
\inf_{\mathrm{diam}\,\mathcal{P}\le\varepsilon} h_\mu(T,\mathcal{P}),
$$
$$
\underline{\mathrm{MRID}}(X,T,\mu,d)
=
\liminf_{\varepsilon\to 0}
\frac{1}{\log(1/\varepsilon)}
\inf_{\mathrm{diam}\,\mathcal{P}\le\varepsilon} h_\mu(T,\mathcal{P}).
$$
These were introduced as a dynamical version of Rényi information dimension through partition entropies [2205.11904], [2010.14772].

A key advantage of the partition formulation is that it yields a variational principle for metric mean dimension **without** tame growth assumptions:
$$
\overline{\mathrm{mdim}}_{\mathrm{M}}(X,T,d)
=
\limsup_{\varepsilon\to 0}
\frac{
\sup_{\mu\in\mathcal{P}_T(X)}
\inf_{\mathrm{diam}\,\mathcal{P}\le\varepsilon} h_\mu(\mathcal{P})
}{
\log(1/\varepsilon)
},
$$
and similarly for the lower metric mean dimension with $\liminf$ [2010.14772]. In this sense, metric mean dimension is the supremum over invariant measures of a mean information-dimension rate defined via shrinking partitions.

The relation between the partition-based and rate-distortion-based formulations is controlled by inequalities such as
$$
R_{\mu,L^\infty}(2\varepsilon)\le
\inf_{\mathrm{diam}\,\mathcal{P}\le\varepsilon} h_\mu(T,\mathcal{P}),
$$
which shows that the partition-based mean Rényi information dimension dominates the $L^\infty$ rate-distortion dimension [2205.11904]. Under the marker property, a further double variational principle holds:
$$
\mathrm{mdim}(X,T)
=
\min_{d\in D(X)}\sup_{\mu\in M(X,T)} \mathrm{MRID}(X,T,\mu,d),
$$
in the sense that the upper and lower versions match in the min–sup identity [2205.11904].

For metrics $d$ that realize mean dimension, namely
$$
D'(X):=\{d\in D(X):\ \mathrm{mdim}(X,T)=\mathrm{mdim}_{\mathrm{M}}(X,T,d)\},
$$
the order of $\sup_\mu$ and $\limsup/\liminf$ can be interchanged under the marker property [2205.11904]. This identifies a regime in which topological and measure-theoretic formulations of mean information dimension coincide exactly.

## 4. Symbolic dynamics and exact entropy formulas

The symbolic setting provides explicit formulas that closely parallel Furstenberg’s one-sided entropy-dimension theorem. Let $A$ be a finite alphabet and $X\subset A^{\mathbb{Z}^2}$ a $\mathbb{Z}^2$-subshift with commuting shifts $\sigma_1,\sigma_2$. For $\alpha>1$, the standard symbolic metric is
$$
d(x,y)=\alpha^{-r(x,y)},
\qquad
r(x,y)=\min\{|u|_\infty: x_u\ne y_u\}.
$$
Under the $\mathbb{Z}$-subaction $T=\sigma_1$, Shinoda and Tsukamoto proved
$$
\mathrm{mdim}_{\mathrm{H}}(X,T,d)
=
\mathrm{mdim}_{\mathrm{M}}(X,T,d)
=
\frac{2\,h_{\mathrm{top}}(X,\sigma_1,\sigma_2)}{\log\alpha},
$$
and for every measure $\mu$ invariant under both $\sigma_1$ and $\sigma_2$,
$$
\mathrm{rdim}(X,\sigma_1,d,\mu)
=
\frac{2\,h_\mu(X,\sigma_1,\sigma_2)}{\log\alpha}.
$$
Hence, for a measure of maximal entropy, mean information dimension, mean Hausdorff dimension, and metric mean dimension all coincide [1910.00844].

The same paper gives directional formulas. For a nonzero vector $(a,b)\in\mathbb{Z}^2$ and $T=\sigma_1^a\sigma_2^b$, one has for the $\ell^\infty$-based symbolic metric
$$
\mathrm{mdim}_{\mathrm{H}}(X,T,d)
=
\mathrm{mdim}_{\mathrm{M}}(X,T,d)
=
\frac{2(|a|+|b|)\,h_{\mathrm{top}}(X,\sigma_1,\sigma_2)}{\log\alpha},
$$
$$
\mathrm{rdim}(X,T,d,\mu)
=
\frac{2(|a|+|b|)\,h_\mu(X,\sigma_1,\sigma_2)}{\log\alpha},
$$
while for the $\ell^2$-radius metric $\rho$ the factor $2(|a|+|b|)$ is replaced by $2\sqrt{a^2+b^2}$ [1910.00844]. The paper interprets these prefactors as capturing the density of orbit traces in the plane in the norm governing the metric.

For the full $\mathbb{Z}^2$ shift $X=A^{\mathbb{Z}^2}$, $h_{\mathrm{top}}(X,\sigma_1,\sigma_2)=\log|A|$, so
$$
\mathrm{mdim}_{\mathrm{M}}=\mathrm{mdim}_{\mathrm{H}}
=
\frac{2\log|A|}{\log\alpha},
$$
and for the Bernoulli product measure,
$$
\mathrm{rdim}(X,\sigma_1,d,\mu)
=
\frac{2\log|A|}{\log\alpha}.
$$
The coefficient $2$ reflects two-sided time in the subaction together with the geometry of the symbolic metric; Furstenberg’s original one-sided result does not have this factor [1910.00844].

These formulas make the phrase “mean information dimension” literal in symbolic dynamics: the number of bits per iterate needed to describe typical orbit segments at distortion $\varepsilon$ scales as an entropy divided by $\log(1/\varepsilon)$.

## 5. Stochastic-process formulation and spectral characterizations

For stationary stochastic processes, mean information dimension is defined through quantized entropy rates. For an $L$-variate process $\{X_t\}$,
$$
d(\{X\})
=
\lim_{m\to\infty}\lim_{k\to\infty}
\frac{H([X^k]_m)}{k\log m},
$$
when the limits exist; equivalently, for stationary processes,
$$
d(\{X\})=\lim_{m\to\infty}\frac{H'(\{[X_t]_m\})}{\log m}.
$$
Geiger and Koch proved that this equals the rate-distortion dimension:
$$
\mathrm{dim}_R(\{X\})=d(\{X\}),
$$
with
$$
\mathrm{dim}_R(\{X\})
=
2\lim_{D\downarrow 0}\lim_{k\to\infty}
\frac{R(X^k,kD)}{-k\log D},
$$
and, for stationary sources,
$$
d(\{X\})=2\lim_{D\downarrow 0}\frac{R(D)}{\log(1/D)}.
$$
Thus the information dimension rate is exactly the high-resolution slope of Shannon’s rate-distortion function [1702.00645].

The same work gives a spectral characterization for stationary Gaussian processes. If $\{X_t\}$ is stationary, $L$-variate, and Gaussian with spectral distribution function $F_X$, then
$$
d(\{X\})
=
\int_{-1/2}^{1/2}\mathrm{rank}\big(F_X'(\theta)\big)\,d\theta.
$$
More generally,
$$
d(\{X\})
\le
\int_{-1/2}^{1/2}\mathrm{rank}\big(F_X'(\theta)\big)\,d\theta,
$$
with equality for Gaussian processes, so among all stationary processes with a given matrix-valued spectral distribution function, the Gaussian process has the largest information dimension rate [1702.00645].

In the scalar Gaussian case, if $S_X(\theta)$ is the power spectral density, then
$$
d(\{X\})
=
\lambda\{\theta\in[-1/2,1/2]: S_X(\theta)>0\},
$$
the Lebesgue measure of the spectral support on which the density is positive. A bandlimited Gaussian process supported on $[-B,B]$ therefore has
$$
d(\{X\})=2B
$$
[1702.00645].

Only the absolutely continuous part of the spectral distribution contributes to the information dimension rate. If the spectral distribution decomposes as
$$
F_X=F_{X,a}+F_{X,d}+F_{X,s},
$$
then $d(\{X\})$ depends only on $F_{X,a}$; the discrete and singular parts contribute zero [1702.00645]. This sharply distinguishes mean information dimension from entropy rate, since a process may carry nontrivial structure in singular or discrete spectral components without increasing the small-distortion dimension rate.

## 6. Extensions, distinctions, and open directions

Several non-equivalent notions of process complexity coexist in the literature. For stationary processes, Geiger and Koch distinguish the information dimension rate $d(\{X\})$ from the block-average notion $d'(\{X\})$ and show that, in general,
$$
d(X_1\mid X^-)\le d(\{X\})\le d'(\{X\}),
$$
with equality under additional mixing hypotheses, such as the condition that there exists $n\ge 0$ with $I(X_k;X_{-n}^- )<\infty$ for all $k$ [1702.00645]. A stationary Gaussian process can satisfy $d(\{X\})<d'(\{X\})$, so “mean information dimension” is not interchangeable with every block-entropy-based dimension notion.

In the shift setting on $[0,1]^{\mathbb{Z}}$, Gutman and Śpiewak showed for ergodic invariant measures that the dynamical mean Rényi information dimension equals the Geiger–Koch information dimension rate [2010.14772]. A later result extended this to all invariant measures on $([0,1]^{\mathbb{Z}},\sigma)$:
$$
\underline{\mathrm{MRID}}([0,1]^{\mathbb{Z}},\sigma,d^{\mathbb{Z}},\mu)=\underline d(\mu),
\qquad
\overline{\mathrm{MRID}}([0,1]^{\mathbb{Z}},\sigma,d^{\mathbb{Z}},\mu)=\overline d(\mu),
$$
and, via the Geiger–Koch theorem, also to the $L^2$ rate-distortion dimensions [2510.08051]. This answers the non-ergodic question left open in the earlier work.

A common misconception is that positive entropy forces positive mean information dimension. Finite-alphabet symbolic systems provide a counterpoint: with the discrete metric, mean Rényi information dimension and rate-distortion dimensions can vanish even though the Kolmogorov–Sinai entropy is positive. At the same time, the associated rate-distortion entropies
$$
h_{\mu,L^p}(T),\quad h_{\mu,L^\infty}(T),\quad h_{\mu,B}(T),\quad \lim_{r\downarrow 0} h_{\mu,r}(T)
$$
equal $h_\mu(T)$ for ergodic measures, and for all invariant measures under the $g$-almost product property [2510.08051]. This indicates that entropy and mean information dimension probe different asymptotic normalizations: fixed-resolution compression versus small-distortion dimensional scaling.

Recent work also sharpens the variational theory. For systems with the marker property and finite mean dimension, the supremum in the double variational principle can be restricted to ergodic measures, and analogous statements hold for a broader family of measure-theoretic $\varepsilon$-entropies, including partition-based, cover-based, Katok, Brin–Katok, and Pfister–Sullivan quantities [2510.08051]. A plausible implication is that the ergodic structure of maximizing measures is more rigid than the original variational formulas alone suggest.

Several open problems remain. Lindenstrauss and Tsukamoto conjectured that the marker property may be unnecessary for the double variational principle [1901.05623]. Yang, Chen, and Zhou likewise isolate the problem of removing the marker-property hypothesis from the interchange of $\sup$ and $\limsup/\liminf$ in the Rényi-information-dimension formulation, and they raise the question of existence of maximal metric mean dimension measures for suitable candidates $F(\mu,\varepsilon)$ [2205.11904]. These questions concern whether mean information dimension can be made fully intrinsic, without auxiliary assumptions on markers or on specially chosen metrics.

Source: https://www.emergentmind.com/topics/mean-information-dimension