Multiscale Complexity Formalism
- Multiscale complexity formalism is a collection of mathematical frameworks that define how information distributions and dependencies vary across scales.
- It integrates axiomatic information theory, coarse-graining, and causal emergence to quantify structure and predict system behavior.
- The approach reveals transitions between localized disorder and global order, aiding in identifying the scales of collective organization.
Searching arXiv for recent and relevant papers on multiscale complexity formalism and closely related frameworks. arXiv search query: "multiscale complexity formalism information-theoretic complexity profile requisite variety emergent complexity" Multiscale complexity formalism denotes a family of mathematical frameworks for characterizing how information, structure, coordination, causal efficacy, or predictive advantage is distributed across scales of description. In the arXiv literature, the phrase covers an axiomatic information-theoretic formalism based on dependencies among components, scale-dependent complexity profiles tied to Ashby’s law of requisite variety, overlap-based and renormalization-based measures for patterns and images, causal-emergence formalisms on partition lattices, and observer-dependent regret-based formulations (Allen et al., 2014, Siegenfeld et al., 2022, Bagrov et al., 2020, Hoel, 17 Mar 2025, Naparstek, 6 Nov 2025).
1. Conceptual scope and the meaning of scale
In one explicitly information-oriented formulation, a scale is “the granularity at which an observer acquires information about a system.” Formally, scale can be represented by a surjective mapping , where is a fine-grained state space and is the coarser state space at scale . A multi-scale system is then a collection of information-processing entities that operate at particular scales, observe other entities at possibly different scales, and may themselves be observed at the same or other scales (Diaconescu et al., 2021).
The same general idea appears in a different computational tradition based on membrane systems. There, a mobile-membrane system uses membrane labels, a rooted tree of nested membranes, and mobility rules such as endocytosis, exocytosis, division, and fusion to encode both within-scale dynamics and inter-scale coupling. The formalism was presented as a “uniform, fully self-contained multiscale framework” in which inter-scale couplings are implemented without external glue functions (Buti et al., 2011).
Taken together, these formulations show that “multiscale complexity formalism” is not a single invariant object. Rather, the literature uses the term for several mathematically explicit ways of relating local and global organization. A plausible implication is that the common denominator is not a unique metric, but the requirement that scale be represented explicitly rather than treated as a purely qualitative intuition.
2. Axiomatic information-theoretic foundations
A central information-theoretic line begins with a set of components and an information function defined for each subset . The axiomatic requirements are monotonicity,
and strong subadditivity,
These conditions are used to build a general theory of structure that is not restricted to Shannon entropy; the same machinery was stated to apply to Shannon entropy, Hartley entropy, Kolmogorov complexity, matroid rank, and vector-space dimension (Allen et al., 2014).
Within this framework, structure is represented by irreducible dependencies among components. Each dependency 0 is assigned an information quantity 1, and a scale 2 equal to the number, or more generally the total weight, of components involved. The complexity profile is then
3
so that 4 measures the information that applies to blocks of size at least 5. The same line of work established an area law,
6
which expresses conservation of total scale-weighted information across scales (Allen et al., 2014).
A related exposition in eco-evolutionary dynamics restated the formalism through an information function on subsets, multivariate mutual information defined by inclusion–exclusion, and a complexity profile 7 built from scale-specific information 8. It also stated additivity under independence: if a system decomposes into two independent subsystems, then their complexity profiles add pointwise (Stacey, 2015).
One recurring point of clarification concerns negative higher-order information. In the parity-bit example on three bits constrained by 9, the three-variable mutual information is negative, 0, even though each pair is independent. In this formalism, that sign is not treated as a paradox but as part of the signed dependency structure of the system (Allen et al., 2014).
3. Complexity profiles, marginal utility, pairwise approximation, and ancillae
A dual summary of multiscale structure is the marginal utility of information. In the descriptor-based formulation, one augments the system by an external descriptor 1 and defines the utility
2
with a budget constraint on the descriptor. The derivative
3
is the Marginal Utility of Information (MUI). MUI curves serve as multiscale complexity profiles in the sense that they quantify how much scale-weighted utility is obtained from an additional bit of description (Stacey, 2017).
Because the full complexity profile is combinatorial in system size, a pairwise approximation was introduced to make the computation tractable. It starts from normalized pairwise mutual informations
4
with 5, and then defines for each variable
6
After inversion and normalization, this yields a pairwise complexity profile computable in 7 time. The pairwise profile preserves linear superposition for unrelated systems and the sum rule, and it is monotonically non-increasing in scale. It agrees with the full profile for the “ideal gas” and “crystal” limits, but it fails on purely higher-order constraints such as the three-bit exclusive-OR system, where all pairwise mutual informations vanish although the triple has a nontrivial constraint (Bar-Yam et al., 2012).
A refinement for more-than-binary variables introduced ancilla components. For each unordered pair of variables,
8
and the ancilla system is the collection of all such 9. The purpose is to expose higher-order structure that is invisible in pairwise mutual informations of the original variables. This construction was used to distinguish James–Crutchfield dyadic and triadic ternary systems, which have the same one-, two-, and three-point Shannon entropies on the original variables but different ancilla distributions and different MUI curves. For the dyadic ancilla, 0 for 1 and 2 thereafter, with 3; for the triadic ancilla, 4 for 5 and 6 thereafter (Stacey, 2017).
These refinements establish an important limitation and an important opportunity. The limitation is that pairwise summaries are not, in general, sufficient. The opportunity is that multiscale formalism can be extended without abandoning the core information-theoretic picture.
4. Scale-dependent complexity and the multi-scale law of requisite variety
A distinct formalism defines complexity directly as a function of scale relative to a nested sequence of partitions. A system 7 of size 8 is modeled as a set of 9 random variables,
0
and an environment for 1 is another system 2 together with a bijection 3. The system 4 “matches” 5 if
6
equivalently 7 (Siegenfeld et al., 2022).
Given a nested sequence of partitions 8, with 9 refining 0, the total information in the parts of 1 is
2
and the discrete complexity profile is the marginal gain
3
This profile satisfies nonnegativity and finite support, and it obeys the sum rule
4
The sum rule is interpreted as a tradeoff between smaller- and larger-scale degrees of freedom: a system cannot be arbitrarily complex at all scales simultaneously (Siegenfeld et al., 2022).
The key theorem is a multi-scale generalization of Ashby’s law. If 5 matches 6, then for every nested partition sequence 7 of 8, the pulled-back partition sequence 9 on 0 satisfies
1
This is the stated multi-scale law of requisite variety. The same work also gives an additivity theorem for independent subsystems,
2
when 3, a replicated-blocks theorem,
4
and a continuum limit obtained from refinements with component size tending to zero (Siegenfeld et al., 2022).
This formalism differs from the earlier dependency-based complexity profile in its choice of primitive object. It does not begin with signed irreducible dependencies, but with partition-relative information increments. The two approaches nonetheless share a conservation principle and a scale-by-scale interpretation.
5. Structural and perceptual variants for patterns, textures, and images
Another major line defines multiscale complexity through coarse-graining and overlap between neighboring renormalized layers. For a pattern 5, one constructs a hierarchy 6 by block-averaging or an RG map. The overlap between successive layers is
7
and the per-scale dissimilarity is
8
The total multiscale complexity is
9
This measure was reported to vanish for both perfectly ordered and purely random patterns, to peak on visually rich structures, and to detect phase boundaries from a single snapshot, with computational cost 0 (Bagrov et al., 2020).
The same structural idea was adapted to visual stimuli as Multi-Scale Structural Complexity (MSSC). There, a multiresolution hierarchy 1 is constructed by low-pass filtering, and the per-scale partial complexity is
2
The total MSSC is
3
For grayscale images computed via FFT-based filtering, the stated computational complexity is 4 (Kravchenko et al., 2024).
Applied to the SAVOIAS dataset, “middle-scale” MSSC yielded the following Pearson correlations with subjective complexity scores:
| Category | 5 |
|---|---|
| Scenes | 0.62 |
| Objects | 0.46 |
| Suprematism | 0.76 |
| Interior design | 0.60 |
| Advertisements | 0.52 |
| Art | 0.36 |
| Infographics | 0.38 |
The same study reported that human judgments rely most strongly on middle spatial bands rather than the finest or coarsest details (Kravchenko et al., 2024).
A related image formalism is the multiscale two-dimensional complexity–entropy causality plane. It treats the spatial lags 6 as the scale, constructs an ordinal-pattern distribution 7, and then computes the normalized permutation entropy
8
and permutation statistical complexity
9
Varying 0 produces a trajectory through the 1 plane and was proposed as a multiscale texture descriptor capable of revealing hidden spatial correlations and roughness crossovers (Zunino et al., 2016).
A common misconception is that randomness is automatically “most complex.” These structural and perceptual formalisms were explicitly motivated by the opposite observation: pure randomness can score highly under entropy or compression criteria while remaining visually less complex than patterns that balance order and disorder (Kravchenko et al., 2024).
6. Causal and observer-dependent generalizations
Recent work recast multiscale complexity in explicitly causal terms. In one formulation, a macroscale is a surjective coarse-graining 2 of a finite Markov chain, and each scale is assigned a causal measure based on interventionist primitives such as sufficiency, necessity, determinism, and specificity. Along a micro-to-macro path 3, the unique contribution at step 4 is
5
Normalizing the positive gains,
6
the emergent complexity is defined as the Shannon entropy
7
In the worked eight-state Markov-chain example, the microscale causal measure was approximately 8, the macro causal measure approximately 9, and the single-step gain approximately 0 bits (Hoel, 17 Mar 2025).
“Engineering Emergence” extended this path-based picture to the full partition lattice. For a partition 1, the macro TPM 2 is assigned determinism and degeneracy, combined into a causal-primitives score
3
and the non-redundant contribution of a scale is
4
Only partitions with 5 belong to the emergent hierarchy. The distribution of gains is then summarized by a path entropy 6 and a row entropy 7, giving a taxonomy of top-heavy, bottom-heavy, mesoscale-peaked, and scale-free/complex hierarchies. The same work described brute-force enumeration as suitable for 8 and a branching-greedy sampling method as scaling to 9 (Jansma et al., 3 Oct 2025).
An observer-dependent alternative is Complexity as Advantage (CAA). For an observer class 00, asymptotic average loss is
01
optimal loss is 02, and regret is 03. Complexity is then defined through the variance of regret across observers or the maximal regret gap. Under log-loss and order-04 Markov observers,
05
and
06
where 07 is excess entropy or predictive information. In this sense, multiscale structure is represented by the profile of predictive gains unlocked by increased observer resources (Naparstek, 6 Nov 2025).
These causal and observer-relative approaches do not replace the earlier information-theoretic formalisms. They shift the emphasis from “how much information exists at a scale” to “what a scale uniquely contributes to causal explanation or prediction.”
7. Representative applications and characteristic regimes
The complexity-profile formalism has been applied directly to two-dimensional ferromagnetic Ising systems. For a finite 08-spin system with Hamiltonian
09
the complexity profile 10 and scale-specific information 11 are defined through conditional mutual informations over subsystems, with the scale-weighted sum rule
12
The reported behavior separates three regimes: in the disordered phase, 13 bits and 14 for 15; near the critical region, a broad spectrum of nonzero 16 appears, 17 has the largest peak amplitude, and large scales acquire non-monotonic structure, including 18 dipping negative; in the ordered phase, 19 bit for every 20 (Al-Azki et al., 29 Jul 2025).
The same study examined pairwise complexity densities 21 and 22 on larger lattices. Both were reported to vanish as 23 and to approach 24 as 25, while each develops a single maximum in the disordered phase below 26; the peak height remains bounded in the thermodynamic limit (Al-Azki et al., 29 Jul 2025).
In eco-evolutionary dynamics, the formalism was used to argue that patterns of interdependency generate structure at multiple scales and that these structures affect evolutionary trajectories. The thesis presentation connected multiscale structure to spatial host–consumer models, multiplayer games, nonlinear payoffs, ecological stochasticity, and generalized conditionalization rules for replicator dynamics in environments with mesoscale structure (Stacey, 2015).
Across these applications, a stable pattern emerges. Ordered systems tend to concentrate information or causal power at large scales, disordered systems concentrate it at the smallest scales, and systems near transitions or under mixed constraints distribute it across several scales. This suggests a shared diagnostic function for multiscale complexity formalisms: they are designed to locate the scales at which collective organization is genuinely present, rather than to assign a single undifferentiated number to an entire system.