Complexity-Entropy Analysis
- Complexity-entropy analysis is a framework that combines entropy measures with complementary complexity notions to quantify disorder and structured organization.
- It employs ordinal-pattern methods, q-complexity-entropy curves, and statistical inference to distinguish between regular, chaotic, and stochastic dynamics.
- Applications range from analyzing solar wind and fractal landscapes to processing images and language, highlighting both methodological strengths and finite-sample challenges.
Complexity-entropy analysis denotes a family of methods that jointly characterize disorder, unpredictability, structural organization, or interaction order by coupling an entropic quantity with a complementary notion of complexity. In current arXiv literature it is not a single formalism but a broad research program spanning ordinal-pattern analysis for time series and images, entropy-rate and excess-entropy representations of symbolic processes, maximum-entropy reconstructions of multivariate dependence, and several physics-motivated constructions in which complexity is treated as an explicit function of entropy or of differences between entropies (Ribeiro et al., 2017, Estevez-Rams et al., 2019, Chliamovitch et al., 2014, Klamut et al., 2020, Datseris et al., 2024).
1. Conceptual scope and competing definitions
No universal definition of complexity-entropy analysis is accepted across the literature. One explicit position is that “complexity” is not universally defined, but in practice often refers to behavior that is neither purely regular nor purely stochastic; many measures called “entropy” or “complexity” differ mainly by representation, probability estimation, and normalization choices rather than by a wholly distinct mathematical content (Datseris et al., 2024).
A major line of work defines complexity through its relation to order and randomness. In the framework based on Gell-Mann’s view, entropy alone is rejected as a universal measure of complexity because both perfect order and complete disorder are treated as simple. The proposed partial measure
therefore vanishes at both entropy extremes and attains its maximum at an intermediate entropy value (Klamut et al., 2020). Closely related in spirit, the two-level-systems literature defines entropic complexity as the difference between Shannon or von Neumann entropy and the second-order Rényi entropy,
so that complexity is again small in the trivial limits of strong localization and complete mixing (Varga, 2024).
A different conceptualization appears in symbolic-sequence analysis. There, complexity is not a single scalar derived from one entropy, but the balance between innovation and context preservation, represented by the entropy rate
and the excess entropy
In this setting, measures irreducible unpredictability and measures useful correlation between past and future (Estevez-Rams et al., 2019).
A more speculative formulation identifies complexity with the time derivative of entropy. Under a deliberately simple information-centered definition, entropy is taken as information content and complexity as “the capacity of a system to incorporate information at a given time,” leading to
That proposal is explicitly presented as a heuristic definitional framework for closed systems rather than as a theorem of statistical mechanics (Modis, 2024).
2. Ordinal-pattern methods and the complexity-entropy plane
The most widely used operational framework is the ordinal-pattern approach of Bandt and Pompe, together with the complexity-entropy causality plane of Rosso and collaborators. For a scalar time series, overlapping segments of length are mapped to ordinal patterns, producing a distribution
The classical normalized permutation entropy is
and the associated statistical complexity is
0
where 1 is the uniform distribution and 2 is a Jensen-Shannon-based disequilibrium (Ribeiro et al., 2017). In the resulting plane, low entropy and low complexity correspond to strong regularity, high entropy and low complexity to noise-like behavior, and intermediate-to-high complexity to structured, patterned dynamics.
A major generalization replaces the single Shannon-based point by an entire parametric 3-complexity-entropy curve. Using Tsallis 4-entropy,
5
and a 6-generalized Jensen-Shannon complexity,
7
the time series is represented by
8
This curve is intended to reveal distinctions that may be weak or invisible at the single point 9, including long-range versus short-range correlations, oscillatory correlations, and chaotic versus stochastic dynamics (Ribeiro et al., 2017).
One of the strongest theoretical results in this literature concerns curve topology. Let 0 be the number of observed ordinal patterns and
1
If all 2 patterns occur, then the 3-curve starts and ends at 4 and is closed. If some patterns are missing, then it begins at 5 and ends at 6, hence it is open (Ribeiro et al., 2017). This open-versus-closed distinction is practically important because deterministic chaotic dynamics often possess forbidden permutations, whereas missing patterns in stochastic series are typically finite-sample effects. The same paper is careful to note that finite stochastic series can also yield open curves, so the topology is diagnostic but not infallible (Ribeiro et al., 2017).
The ordinal framework extends naturally to two-dimensional data. For a 7 sliding window over an image, the number of local ordinal states is
8
the normalized permutation entropy is
9
and statistical complexity is
0
with 1 the uniform distribution (Ribeiro et al., 2012). This extension has been applied to fractal landscapes, liquid-crystal textures, and Ising surfaces, where it identified roughness changes, phase transitions, and critical temperature, respectively (Ribeiro et al., 2012).
3. Statistical inference and computational infrastructure
Recent work has moved ordinal complexity-entropy analysis from descriptive plotting toward asymptotic inference. Under weak dependence conditions, the empirical ordinal-pattern frequency vector satisfies
2
and this central limit behavior propagates to entropy and complexity (Silbernagel et al., 23 Jul 2025). The key distinction is whether the ordinal-pattern distribution is uniform.
In the non-uniform case 3, the first-order delta method applies and the normalized entropy-complexity pair is asymptotically bivariate normal. In the uniform case 4, the entropy gradient vanishes, first-order arguments fail, and the limit law becomes a quadratic form. Asymptotically, the pair lies on a straight line through 5 in the entropy-complexity plane, with line
6
in the notation of the paper (Silbernagel et al., 23 Jul 2025). A further consequence is that entropy and complexity are often extremely highly negatively correlated, so complexity may add little inferential power beyond entropy for testing serial dependence (Silbernagel et al., 23 Jul 2025).
The methodological proliferation of entropy and complexity measures has also led to software unification efforts. A notable example is “ComplexityMeasures.jl,” which presents complexity-entropy analysis as a composable pipeline: define an outcome space, map data into outcomes, estimate probabilities, and apply an information or complexity functional. The package reports 1638 measures with 3,841 lines of source code in version 3.7, and organizes PMF-based analysis around calls such as 9 together with a separate 0 interface for non-PMF-based measures (Datseris et al., 2024). This software perspective emphasizes that many named measures differ only in symbolization, probability estimation, or the functional applied to the resulting distribution.
4. Non-ordinal formulations
Complexity-entropy analysis also appears in frameworks that do not rely on ordinal patterns. In written language, one established formulation places texts in an 7-versus-8 diagram, where entropy rate measures innovation and excess entropy measures context preservation. Lempel-Ziv estimates are used for both quantities, and controlled randomizations of sentences, words, or characters isolate the contributions of different organizational levels (Estevez-Rams et al., 2019). In that literature, complexity is not identified with a separate statistical complexity functional, but with the joint position in the 9-0 plane.
A different line uses maximum-entropy reconstruction to decompose statistical dependence by interaction order. If 1 denotes the maximum-entropy approximation constrained by marginals up to order 2, then the incremental contribution of order 3 is
4
where 5 is Kullback-Leibler divergence. Summing these terms yields the multi-information
6
Complexity, in this sense, is the extent to which low-order information fails to reconstruct the full distribution, so high-order dependencies are essential (Chliamovitch et al., 2014).
Quantum and statistical-physics applications often employ composite measures rather than planes. For the 7-dimensional rigid rotator, explicit entropic moments and Rényi entropies are combined with Fisher information to define Fisher-Rényi, Fisher-Shannon, and LMC complexities, and Fisher-Shannon is reported to follow the intuitive angular lobe structure most faithfully (Dehesa et al., 2015). For hydrogenic Rydberg atoms, the same families—Cramér-Rao, Fisher-Shannon, and LMC—are studied in both position and momentum space, with asymptotic growth laws derived from Laguerre and Gegenbauer polynomial asymptotics (López-Rosa et al., 2013). In two-level systems, the entropy gap 8 plays the role of a basis-insensitive structural complexity and becomes maximal when ordering and disordering mechanisms compete (Varga, 2024).
5. Representative application domains
Complexity-entropy analysis has been applied across nonlinear dynamics, image analysis, plasma physics, language, and cosmological or historical modeling. The following cases illustrate the diversity of interpretations.
| Domain | Setting | Reported pattern |
|---|---|---|
| Solar wind | Magnetic-field fluctuations at 1 au | Fast wind has the highest entropy and lowest complexity; magnetic clouds have the lowest entropy and highest complexity; differences sharpen with timescale (Kilpua et al., 2024) |
| Solar photosphere | Hinode images of 9 and horizontal electromagnetic energy flux | During a 37.5 min vortex expansion, complexity rises and entropy falls, consistent with coherent-structure formation and an inverse turbulent cascade (Chian et al., 22 Sep 2025) |
| Fractal, stochastic, and chaotic signals | Weierstrass function, colored noise, logistic map | Complexity rises with Weierstrass fractional dimension; 1/f noise has the highest complexity; logistic-map entropy maps follow the bifurcation diagram (Brechtl et al., 2017) |
| Written language | English texts with surrogate randomizations | Authors occupy distinct regions in 0-1 space; the largest organizational contribution comes from letters forming words (Estevez-Rams et al., 2019) |
| Images and textures | Fractal landscapes, liquid crystals, Ising surfaces | The 2D complexity-entropy plane distinguishes textures, identifies liquid-crystal transitions, and detects the Ising critical temperature (Ribeiro et al., 2012) |
These applications show that the same formal vocabulary—entropy, complexity, disequilibrium, forbidden patterns, excess entropy, or multiscale entropy—can serve rather different scientific purposes. In solar-wind analysis, the central distinction is between turbulence-dominated and coherent large-scale structures; at small scales the different wind types look similar, while at larger scales magnetic clouds move toward lower entropy and higher complexity (Kilpua et al., 2024). In photospheric turbulence, a trajectory in the plane is used to quantify a transition from fragmented inhomogeneity to organized coherence, with the authors interpreting the resulting location as an admixture of chaos and stochasticity rather than a purely random state (Chian et al., 22 Sep 2025).
Language studies use the framework to partition organization across letters, words, and sentences rather than to classify chaos. The excess-entropy/entropy-rate representation places Abbott, Doyle, and Shakespeare in distinct regions, and controlled scrambling demonstrates that sentence shuffling raises entropy rate and lowers excess entropy, while character shuffling drives excess entropy nearly to zero (Estevez-Rams et al., 2019). In contrast, the natural-versus-artificial-language study based on normalized Shannon entropy, emergence, self-organization, and 2 treats complexity as a deterministic function of entropy and applies it to English, Spanish, and software code (Febres et al., 2013).
6. Limitations, controversies, and current directions
The first limitation is definitional. Some papers explicitly dispute the view that larger entropy means larger complexity and instead insist that complexity should peak between full order and full disorder (Klamut et al., 2020). Others nevertheless use entropy alone as a proxy for complexity in language analysis, interpreting lower entropy as higher complexity (Xie et al., 2016). The heuristic proposal 3 provides yet another definition-driven relation, but it is explicitly not a rigorous derivation from thermodynamics or information theory (Modis, 2024). Comparison across studies is therefore nontrivial even when the same vocabulary is used.
A second limitation is dependence on representation and parameter choice. Ordinal methods depend on embedding dimension, delay, window size, and the adequacy of estimating a distribution over 4 or 5 states. The time-series literature repeatedly stresses the requirement 6, and the solar-wind application restricts the maximum lag to maintain robustness criteria (Ribeiro et al., 2017, Kilpua et al., 2024). Software-oriented work likewise emphasizes that changing outcome space, estimator, or stencil can materially shift the location of a system in the entropy-complexity plane (Datseris et al., 2024).
A third limitation concerns inference and finite-sample ambiguity. Open 7-complexity-entropy curves often suggest forbidden patterns and deterministic constraints, but finite stochastic series can also yield open curves that close only as 8 increases (Ribeiro et al., 2017). The asymptotic theory of ordinal entropy-complexity pairs shows that the correct limit law depends sharply on whether the ordinal distribution is uniform, so naive Gaussian uncertainty estimates can be invalid near the i.i.d. case (Silbernagel et al., 23 Jul 2025).
Finally, some formulations are intentionally speculative. The cosmological-historical proposal that complexity behaves like the derivative of entropy applies closed-system reasoning to an anthropic, nonequilibrium, open historical sequence of 28 milestones and acknowledges substantial subjectivity in event selection, equal-importance assumptions, and forecasting (Modis, 2024). A plausible implication is that the field is converging not toward one canonical complexity measure, but toward a layered methodology: explicitly state the entropy notion, explicitly state the complexity notion, justify the representation, and quantify uncertainty whenever possible.