---
title: Entropy-Complexity Pair Overview
url: https://www.emergentmind.com/topics/entropy-complexity-pair
type: topic
---

# Entropy-Complexity Pair Overview

An entropy-complexity pair is a joint characterization in which an entropy quantity is considered together with a complexity quantity, so that randomness, disorder, or logarithmic class size is not treated in isolation from structure, predictability, disequilibrium, or description cost. Across the cited literature, the pair appears in multiple, non-equivalent forms: as a point or curve in a complexity-entropy plane for time series, as the relation between excess entropy and statistical complexity in computational mechanics, as a correspondence between Boltzmann entropy and formula-length complexity in logic, as an equality between entropy rate and algorithmic complexity rate in ergodic dynamics, and as a Kähler-geometric pairing for coherent states [1705.04779] [0905.2918] [2209.12564] [2407.13327].

## 1. Conceptual range of the pair

The recurring pattern is that entropy quantifies uncertainty, disorder, or class size, whereas complexity quantifies structure, stored predictive information, disequilibrium, description length, or state-to-state separation. The literature therefore uses the same phrase for mathematically different objects. This suggests that “entropy-complexity pair” is best understood as a family resemblance across frameworks rather than a single invariant construction.

| Domain | Entropy quantity | Complexity quantity |
|---|---|---|
| Ordinal time-series analysis | \(H_1(P)\), \(H_q(P)\), \(H_\alpha(P)\) | \(C_1(P)\), \(C_q(P)\), \(C_\alpha(P)\) |
| Computational mechanics | \(E=I[\overleftarrow{X};\overrightarrow{X}]\) | \(C_\mu=H[\mathcal{S}_0]\) |
| Logic-based model classes | \(S_B(C_i)=\log_2|C_i|\) | \(DC_L(M)\) |
| Coherent-state geometry | Shannon entropy \(S\) | \(\ln C(z,w)=D(z,w)\) |

In the time-series setting, the entropy-complexity pair usually combines a normalized entropy with a disequilibrium factor relative to the uniform distribution, producing either a single point or a parametric curve [1705.04779] [1801.05738]. In computational mechanics, the pair \(E\) and \(C_\mu\) separates predictive information stored in the process from memory required by a predictive model, and their difference is interpreted as information erasure [0905.2918]. In the logic-based setting, Boltzmann entropy is the logarithm of the size of an equivalence class of structures, whereas description complexity is the minimum formula size needed to define that class [2209.12564]. In coherent-state geometry, entropy and complexity both descend from the same Kähler potential of the Fubini-Study metric [2407.13327].

A recurrent misconception is that entropy itself is a direct measure of complexity. One paper explicitly disputes this view and introduces a nonlinear transformation of time-dependent entropy,
\[
CX(S;m,n)=(S^{\max}-S)^m(S-S^{\min})^n,
\]
with \(CX\to0\) at both \(S^{\min}\) and \(S^{\max}\), and a unique maximum at
\[
S_{CX}^{\max}=\frac{mS^{\min}+nS^{\max}}{m+n}.
\]
In that formulation, the most complex states are “optimally mixed” rather than maximally entropic [2006.01900].

## 2. Ordinal-pattern time series and complexity-entropy curves

In the complexity-entropy plane for time series, one begins with a probability distribution \(P=\{p_1,\dots,p_{d!}\}\) over ordinal patterns. For Shannon entropy,
\[
S_1(P)=-\sum_{i=1}^{d!} p_i\log p_i,\qquad
H_1(P)=\frac{S_1(P)}{\log(d!)}.
\]
The corresponding classical statistical complexity is
\[
C_1(P)=\frac{D_1(P,U)\,H_1(P)}{D_1^*},
\]
where \(U=\{1/d!\}\) is the uniform distribution, \(D_1(P,U)\) is the Jensen-Shannon divergence,
\[
D_1(P,U)=S_1\!\left(\frac{P+U}{2}\right)-\frac{S_1(P)+S_1(U)}{2},
\]
and \(D_1^*\) is its maximum over all \(P\) [1705.04779].

The Tsallis generalization replaces Shannon entropy by
\[
S_q(P)=\sum_{i=1}^{d!} p_i\log_q(1/p_i),\qquad
\log_q(x)=\frac{x^{1-q}-1}{1-q},
\]
with normalized entropy
\[
H_q(P)=\frac{S_q(P)}{\log_q(d!)}.
\]
The corresponding \(q\)-complexity is
\[
C_q(P)=\frac{D_q(P,U)\,H_q(P)}{D_q^*},
\]
where
\[
D_q(P,U)=\frac12 K_q\!\left(P\middle\|\frac{P+U}{2}\right)+\frac12 K_q\!\left(U\middle\|\frac{P+U}{2}\right),
\]
and
\[
K_q(P\|R)=-\sum_i p_i\log_q(r_i/p_i).
\]
The \(q\)-complexity-entropy curve is the parametric set \(\{(H_q(P),C_q(P))\}\) as \(q\) varies [1705.04779].

The distribution \(P\) is typically estimated by the Bandt-Pompe method. Given a scalar time series \(\{x_1,\dots,x_n\}\) and embedding dimension \(d\), one forms the overlapping vectors \(v_s=(x_{s-(d-1)},\dots,x_s)\), determines the permutation that sorts each vector in ascending order, and counts how many times each of the \(d!\) possible permutations occurs. The empirical frequencies
\[
p_j=\frac{\#\text{ occurrences of }\pi_j}{n-d+1}
\]
capture the “ordinal dynamics” at scale \(d\) [1705.04779].

The interpretation of the resulting curves is highly specific. Stochastic processes such as white noise, fractional Brownian motion, and harmonic noise eventually sample all \(d!\) patterns, so \(P\) has full support and the \(q\)-curve is a closed loop starting and ending at \((H_q\to1,C_q\to0)\) as \(q\to0\) or \(q\to\infty\). Chaotic maps often forbid some ordinal patterns, so \(P\) has zeros and the curve is open. Long-range correlations produce broader loops; short-range or uncorrelated signals produce narrower loops; oscillatory correlations distort loop width as a function of \(q\). The extremal values \(q_H^*\), where \(H_q(P)\) is minimal, and \(q_C^*\), where \(C_q(P)\) is maximal, serve as discriminating scales. In the reported applications, the curves distinguish simulated stochastic processes and chaotic maps, yield open curves for chaotic laser-intensity pulsations, show closed or partly open loops for crude-oil prices and the sunspot index, and improve heart-rate classification from approximately \(76\%\) at \(q=1\) to approximately \(80\%\) when \((H_{q_H^*},C_{q_C^*})\) is used in a k-NN classifier [1705.04779].

A parallel construction uses Rényi entropy,
\[
S_\alpha(P)=\frac{1}{1-\alpha}\ln\!\left(\sum_{i=1}^n p_i^\alpha\right),\qquad
H_\alpha(P)=\frac{S_\alpha(P)}{\ln n},
\]
together with a Rényi-based disequilibrium \(D_\alpha(P)\) and normalized complexity
\[
C_\alpha(P)=\frac{D_\alpha(P)\,H_\alpha(P)}{D_\alpha^*}.
\]
As \(\alpha\) varies, \((H_\alpha,C_\alpha)\) traces a Rényi complexity-entropy curve. Its local geometry near \(\alpha\downarrow0\) is used for classification: stochastic processes yield positive curvature near \((1,0)\), chaotic processes yield negative curvature from an initial point inside the unit square when forbidden patterns are present, and periodic time series produce vertical straight lines because \(H_\alpha\) is independent of \(\alpha\) when the realized patterns are equally likely [1801.05738].

## 3. Asymptotic theory and statistical inference

The entropy-complexity pair became a “popular tool for summarizing the time-series dynamics,” but its inferential theory was developed later. For ordinal patterns of length \(m\), let \(p=(p_1,\dots,p_{m!})^\top\) be the true ordinal-pattern distribution and
\[
\hat p_i=\frac1n\sum_{t=1}^n \mathbf 1\{\Pi_t=\pi_i\}
\]
its empirical estimator. Under weak dependence conditions, one has
\[
\sqrt n\,(\hat p-p)\xrightarrow{d}N(0,\Sigma),
\]
where the long-run covariance matrix has entries
\[
\Sigma_{ij}=p_i(\delta_{ij}-p_j)+\sum_{k=1}^\infty \bigl(p_{ij}(k)+p_{ji}(k)-2p_ip_j\bigr).
\]
From this, a delta-method analysis yields the asymptotic distribution of the entropy-complexity pair [2507.17625].

A key distinction is between the non-uniform case \(p\neq u\) and the uniform case \(p=u\). In the non-uniform case, the first-order delta method applies to the smooth map \(\Psi(p)=(H_0H(p),C(p))^\top\), giving
\[
\sqrt n\bigl(\Psi(\hat p)-\Psi(p)\bigr)\xrightarrow{d}N(0,\Sigma^{(3)}).
\]
In the uniform case, the first-order delta method collapses, and a second-order Taylor expansion of entropy is required. Then
\[
n\bigl(H(\hat p)-\log(m!)\bigr)\dto -\frac{m!}{2}Q_m,
\]
where \(Q_m\) is a quadratic form determined by the nonzero eigenvalues of \(\Sigma\). A corresponding result holds jointly for \((H_0H(\hat p),C(\hat p))\) [2507.17625].

These limit theorems support formal inference. Under the i.i.d. null, the ordinal-pattern distribution is uniform, and the pair concentrates on the line
\[
C=\frac{D_0}{\log(m!)}(1-H).
\]
Two one-dimensional tests were proposed,
\[
\widehat{HD}=\frac{n}{m!}\bigl(\log(m!)-H(\hat p)+4D(\hat p)\bigr),\qquad
\widehat{HC}=\frac{n}{m!}\bigl(\log(m!)-H(\hat p)+\tfrac4{D_0}C(\hat p)\bigr),
\]
both converging to \(Q_m\) under the null. For \(m=3\), \(\alpha=0.05\), and sample sizes \(T=100,250,500,1000\), the reported finite-sample results show size approximately \(0.05\) for i.i.d. \(N(0,1)\), power tending to \(1\) for AR(1), and moderate power for QMA(1) and TEAR(1). In those experiments, complexity added no appreciable power gain beyond entropy alone. The same asymptotic theory also yields approximate standard errors and confidence ellipsoids for \((H_0H(\hat p),C(\hat p))\) when \(p\neq u\) [2507.17625].

## 4. Predictive information, statistical complexity, and criticality

In computational mechanics, the canonical entropy-complexity pair is excess entropy and statistical complexity. For a stationary stochastic process with infinite past \(\overleftarrow{X}\) and infinite future \(\overrightarrow{X}\), excess entropy is
\[
E=I[\overleftarrow{X};\overrightarrow{X}]
=\sum_{L=1}^\infty (h_L-h_\mu),
\]
where \(h_L=H[L]-H[L-1]\) and \(h_\mu=\lim_{L\to\infty}h_L\). Causal states \(\mathcal S\) partition pasts according to equality of future conditional distributions, and statistical complexity is
\[
C_\mu=H[\mathcal S_0].
\]
The central relation is
\[
\Delta\equiv C_\mu-E=H[\mathcal S_0\mid \overrightarrow{X}],
\]
so the gap between stored memory and predictive information is exactly the information erased during forecasting. The efficiency ratio
\[
\eta=\frac{E}{C_\mu}=1-\frac{\Delta}{C_\mu}
\]
satisfies \(0\le \eta\le1\). In this framework, models with small \(\Delta\) or large \(\eta\) are the efficient ones, and \(\Delta=0\) corresponds to “crypticity zero” [0905.2918].

For finite symbolic data, one may use an approximate complexity proxy. In the analysis of the single-spin time series of the \(2D\) Ising ferromagnet, the \(M\)-block Shannon entropy is
\[
H_M=-\sum_{s^M\in\{0,1\}^M} P(s^M)\log_2 P(s^M),
\]
and the entropy rate is
\[
h=\lim_{M\to\infty}[H_M-H_{M-1}].
\]
An “approximate complexity” is then defined by
\[
c_{\rm approx}=1-\frac{h}{h_1},
\]
with \(h_1=H_1\), or operationally by estimating \(h_1\) from a shuffled sequence that preserves symbol frequencies but destroys temporal correlations. Three estimators were used: block entropy, NSRPS, and zlib/LZ77 compression [1206.7032].

When the spin orientation of a fixed site is recorded under Metropolis dynamics on a \(256\times256\) lattice, the entropy-complexity pair exhibits the characteristic critical pattern. The entropy rate \(h(T)\) grows from \(0\) at \(T\to0\), passes through an inflection near \(T_c\approx2.269\), and saturates at \(1\) for \(T\gg T_c\). The approximate complexity \(c_{\rm approx}(T)\) is small in both the ordered and disordered phases but has a pronounced peak at \(T\approx T_c\). The parametric diagram \(c_{\rm approx}\) versus \(h\) forms a loop rising from \((0,0)\), turning back near \((h(T_c),c_{\rm approx}(T_c))\), and returning to \((1,0)\). A finite-size scaling fit for the peak location,
\[
T_{\rm peak}(N)=T_c+aN^{-b},
\]
gave
\[
T_{\rm peak}(N)=2.270(5)+0.9\,N^{-0.50(2)},
\]
showing convergence of the complexity maximum to the true critical temperature [1206.7032].

## 5. Algorithmic, logical, and dynamical correspondences

A different use of the entropy-complexity pair arises when entropy rate is compared with algorithmic or description complexity rate. In the classical dynamical setting, Brudno’s theorem states that for an ergodic system the Kolmogorov complexity rate of almost every symbolic trajectory equals the Kolmogorov-Sinai entropy rate. In the notation of prefix-free Kolmogorov complexity \(C(x)\),
\[
\limsup_{n\to\infty}\frac1n C(i_{1:n})=h_{KS}(T)
\]
for \(\mu\)-almost every orbit. Using Gács’ machine-independent complexity, the same equality is recovered in the classical ergodic case, and a quantum analogue is proved densely for ergodic quantum spin chains:
\[
\lim_{n\to\infty}\frac{1}{2n+1}\,\widehat H(\pi)=s(\omega)
\]
for minimal projectors inside the typical subspaces provided by the quantum Shannon-McMillan theorem [1705.09449].

The amenable-group generalization extends this correspondence beyond \(\mathbb Z\)-actions. For a computable amenable group \(G\), a modest, tempered, computable Følner sequence \(\{F_n\}\), and an ergodic invariant measure \(\mu\) on \(\Sigma^G\),
\[
\lim_{n\to\infty}\frac{K(x|_{F_n})}{|F_n|}=h_\mu(\sigma)
\]
for \(\mu\)-almost every \(x\). In the topological case, every point satisfies
\[
\overline K(x)\le h_{\mathrm{top}}(\sigma,X),
\]
and there exists \(x^*\) attaining equality. This makes entropy a precise rate of algorithmic information production for typical or entropy-maximizing orbits of amenable group actions [1809.01634].

In the logic-based scenario, the entropy side is Boltzmann entropy of equivalence classes,
\[
S_B(C_i)=\log_2|C_i|,
\]
and the complexity side is minimal formula length,
\[
DC_L(M)=\min\{\mathrm{size}(\varphi)\mid \varphi\in L,\ \varphi\ \text{defines}\ M\}.
\]
For MLU over a finite unary vocabulary, the class realizing all \(1\)-types has maximal Boltzmann entropy and also maximal description complexity. For GMLU,
\[
\langle H_B\rangle = k\,n+o(n),\qquad \langle C\rangle=n+o(n),
\]
so
\[
\langle H_B\rangle \sim k\,\langle C\rangle.
\]
By contrast, for first-order logic over vocabularies with at least one relation of arity \(m\ge2\), expected description complexity grows asymptotically faster than expected Boltzmann entropy. The paper therefore shows that the entropy-complexity correspondence depends sharply on the expressive power of the description language [2209.12564].

## 6. Geometric and thermodynamic realizations

For spin-\(\tfrac12\) coherent states of \(SL(2,\mathbb C)\) on \(\mathbb{CP}^1\), entropy and complexity both descend from the Fubini-Study Kähler potential
\[
K(z,\bar z)=\ln(1+|z|^2).
\]
The coherent-state probabilities are
\[
p_0=\frac1{1+|z|^2},\qquad p_1=\frac{|z|^2}{1+|z|^2},
\]
with Shannon entropy
\[
S=\ln(1+|z|^2)-\frac{|z|^2}{1+|z|^2}\ln|z|^2.
\]
After Legendre transforming to the symplectic coordinate \(y=1/(1+|z|^2)\), the Guillemin potential becomes
\[
g(y)=y\ln y+(1-y)\ln(1-y),
\]
and one has
\[
S=-g(y).
\]
For two coherent states \(|z\rangle\) and \(|w\rangle\), complexity is defined by
\[
C(z,w)=\frac{1}{|\langle w|z\rangle|^2},
\]
so that
\[
\ln C(z,w)=D(z,w),
\]
where \(D\) is Calabi’s diastasis. In this construction, entropy and log-complexity are two different manifestations of the same Kähler potential. Non-trivial deformations of the Kähler potential break the precise equality \(\ln C=D\), which is why the Fubini-Study metric is described there as “optimal” [2407.13327].

Thermodynamic versions of the pair link entropy or sequence complexity to extractable work. In the Mandal-Jarzynski model, a two-state system interacts with a tape of bits, and the average work obeys
\[
\Delta W\le k_BT\,\Delta H,
\]
where \(\Delta H\) is the Shannon-entropy increase of the tape. The same argument extends to any concave entropy-like functional \(\bar h(p)\), including the predictability functional \(\bar h(p)=\min\{p,1-p\}\) and the squared-error functional \(\bar h(p)=p(1-p)\). For an individual incoming bit-string \(x^n\), Shannon entropy can be replaced by the Lempel-Ziv complexity
\[
\rho(x^n)=\frac{c(x^n)\ln c(x^n)}{n},
\]
yielding, in the slow-mixing limit,
\[
\langle W_n\rangle\le k_BT\,n[\ln2-\rho(x^n)].
\]
This gives LZ complexity the role of an individual-sequence entropy in a physical work-extraction bound [1503.07653].

A related classical Markov-chain model of \(K\) independent \(n\)-dits defines configuration entropy by
\[
S(a_0,\dots,a_{n-1})=\ln\frac{K!}{\prod_i a_i!},
\]
and classical absolute complexity by the minimum of the complexity relative to all-uniform reference states. In that model, states of maximal entropy coincide with states of maximal absolute complexity, and a “Second Law of Classical Absolute Complexity” is conjectured: in an isolated system, \(C_{\rm abs}(t)\) tends to grow on average until it reaches its maximum. For bits, the average complexity satisfies
\[
C(t)=\frac K2\left[1-\left(1-\frac2K\right)^t\right],
\]
which defines a parametric trajectory in the \((C,S)\)-plane [1902.10538].

Specialized physical applications also use paired information-theoretic measures. For hydrogenic Rydberg atoms, Shannon entropy and Fisher information are combined with Cramér-Rao, Fisher-Shannon, and López-Mancini-Calbet complexities in both position and momentum space. For large principal quantum number \(n\), the Cramér-Rao and Fisher-Shannon complexities grow as \(\sim n^2\) for several families of states, whereas some LMC complexities saturate to \(O(1)\) constants, such as \(C_{LMC}[r]\to2\) and \(C_{LMC}[p]\to10\) for circular states [1305.1149].

Taken together, these formulations show that entropy-complexity pairs are not interchangeable. Some pairs diagnose forbidden ordinal patterns, some quantify predictive inefficiency, some express equality theorems between entropy rate and algorithmic complexity rate, and some derive both quantities from a single geometric or thermodynamic structure. A plausible implication is that the pair is most informative when the chosen complexity notion is matched to the operative structure of the problem: ordinal support for time series, causal states for prediction, formula size for definability, or Kähler potential for coherent-state geometry.

Source: https://www.emergentmind.com/topics/entropy-complexity-pair