---
title: Multiscale Complexity Formalism
url: https://www.emergentmind.com/topics/multiscale-complexity-formalism
type: topic
---

# Multiscale Complexity Formalism

Searching arXiv for recent and relevant papers on multiscale complexity formalism and closely related frameworks.
arXiv search query: "multiscale complexity formalism information-theoretic complexity profile requisite variety emergent complexity"
Multiscale complexity formalism denotes a family of mathematical frameworks for characterizing how information, structure, coordination, causal efficacy, or predictive advantage is distributed across scales of description. In the arXiv literature, the phrase covers an axiomatic information-theoretic formalism based on dependencies among components, scale-dependent complexity profiles tied to Ashby’s law of requisite variety, overlap-based and renormalization-based measures for patterns and images, causal-emergence formalisms on partition lattices, and observer-dependent regret-based formulations [1409.4708, 2206.04896, 2003.04632, 2503.13395, 2511.04590].

## 1. Conceptual scope and the meaning of scale

In one explicitly information-oriented formulation, a scale is “the granularity at which an observer acquires information about a system.” Formally, scale can be represented by a surjective mapping $\phi_\sigma:S\to A_\sigma$, where $S$ is a fine-grained state space and $A_\sigma$ is the coarser state space at scale $\sigma$. A multi-scale system is then a collection of information-processing entities that operate at particular scales, observe other entities at possibly different scales, and may themselves be observed at the same or other scales [2106.01801].

The same general idea appears in a different computational tradition based on membrane systems. There, a mobile-membrane system $M=(\Omega,\Gamma,\mu,w_0,\mathcal R,\Lambda,\sigma)$ uses membrane labels, a rooted tree of nested membranes, and mobility rules such as endocytosis, exocytosis, division, and fusion to encode both within-scale dynamics and inter-scale coupling. The formalism was presented as a “uniform, fully self-contained multiscale framework” in which inter-scale couplings are implemented without external glue functions [1108.3434].

Taken together, these formulations show that “multiscale complexity formalism” is not a single invariant object. Rather, the literature uses the term for several mathematically explicit ways of relating local and global organization. A plausible implication is that the common denominator is not a unique metric, but the requirement that scale be represented explicitly rather than treated as a purely qualitative intuition.

## 2. Axiomatic information-theoretic foundations

A central information-theoretic line begins with a set of components $A$ and an information function $H(U)$ defined for each subset $U\subset A$. The axiomatic requirements are monotonicity,
$$
H(U)\le H(V)\quad\text{if }U\subset V\subset A,
$$
and strong subadditivity,
$$
H(U\cup V)\le H(U)+H(V)-H(U\cap V).
$$
These conditions are used to build a general theory of structure that is not restricted to Shannon entropy; the same machinery was stated to apply to Shannon entropy, Hartley entropy, Kolmogorov complexity, matroid rank, and vector-space dimension [1409.4708].

Within this framework, structure is represented by irreducible dependencies among components. Each dependency $x$ is assigned an information quantity $I(x)$, and a scale $s(x)$ equal to the number, or more generally the total weight, of components involved. The complexity profile is then
$$
C_A(k)=\sum_{\substack{x\in D_A\\ s(x)\ge k}} I(x),
$$
so that $C_A(k)$ measures the information that applies to blocks of size at least $k$. The same line of work established an area law,
$$
\sum_{k\ge1} C_A(k)=\sum_{a\in A} H(a),
$$
which expresses conservation of total scale-weighted information across scales [1409.4708].

A related exposition in eco-evolutionary dynamics restated the formalism through an information function on subsets, multivariate mutual information defined by inclusion–exclusion, and a complexity profile $C(k)$ built from scale-specific information $D(\ell)=C(\ell)-C(\ell+1)$. It also stated additivity under independence: if a system decomposes into two independent subsystems, then their complexity profiles add pointwise [1509.02958].

One recurring point of clarification concerns negative higher-order information. In the parity-bit example on three bits constrained by $a\oplus b\oplus c=0$, the three-variable mutual information is negative, $I(a;b;c)=-1$, even though each pair is independent. In this formalism, that sign is not treated as a paradox but as part of the signed dependency structure of the system [1409.4708].

## 3. Complexity profiles, marginal utility, pairwise approximation, and ancillae

A dual summary of multiscale structure is the marginal utility of information. In the descriptor-based formulation, one augments the system by an external descriptor $d$ and defines the utility
$$
U(d)=\sum_{s\in S}\sigma(s)\,I(d;s),
$$
with a budget constraint on the descriptor. The derivative
$$
M(y)=\frac{d}{dy}\bigl[\max_{H'(d)=y}U(d)\bigr]
$$
is the Marginal Utility of Information (MUI). MUI curves serve as multiscale complexity profiles in the sense that they quantify how much scale-weighted utility is obtained from an additional bit of description [1705.03927].

Because the full complexity profile is combinatorial in system size, a pairwise approximation was introduced to make the computation tractable. It starts from normalized pairwise mutual informations
$$
A_{ij}=I(X_i;X_j),\qquad 0\le A_{ij}\le1,
$$
with $A_{ii}=1$, and then defines for each variable
$$
k_i(p)=\bigl|\{\,j\mid A_{ij}\ge p\}\bigr|.
$$
After inversion and normalization, this yields a pairwise complexity profile computable in $O(N^2\log N)$ time. The pairwise profile preserves linear superposition for unrelated systems and the sum rule, and it is monotonically non-increasing in scale. It agrees with the full profile for the “ideal gas” and “crystal” limits, but it fails on purely higher-order constraints such as the three-bit exclusive-OR system, where all pairwise mutual informations vanish although the triple has a nontrivial constraint [1208.0823].

A refinement for more-than-binary variables introduced ancilla components. For each unordered pair of variables,
$$
A_{ij}=
\begin{cases}
1 & \text{if } a_i=a_j,\\
0 & \text{otherwise,}
\end{cases}
$$
and the ancilla system is the collection of all such $A_{ij}$. The purpose is to expose higher-order structure that is invisible in pairwise mutual informations of the original variables. This construction was used to distinguish James–Crutchfield dyadic and triadic ternary systems, which have the same one-, two-, and three-point Shannon entropies on the original variables but different ancilla distributions and different MUI curves. For the dyadic ancilla, $M_D(y)=3$ for $0\le y\le h$ and $0$ thereafter, with $h\approx0.8113$; for the triadic ancilla, $M_T(y)=2$ for $0\le y\le2$ and $0$ thereafter [1705.03927].

These refinements establish an important limitation and an important opportunity. The limitation is that pairwise summaries are not, in general, sufficient. The opportunity is that multiscale formalism can be extended without abandoning the core information-theoretic picture.

## 4. Scale-dependent complexity and the multi-scale law of requisite variety

A distinct formalism defines complexity directly as a function of scale relative to a nested sequence of partitions. A system $X$ of size $N$ is modeled as a set of $N$ random variables,
$$
X=\{x_1,\dots,x_N\},
$$
and an environment for $X$ is another system $Y=\{y_1,\dots,y_N\}$ together with a bijection $f:Y\to X$. The system $X$ “matches” $(Y,f)$ if
$$
H(y_i\mid x_i)=0\quad\forall i,
$$
equivalently $H(Y\mid X)=0$ [2206.04896].

Given a nested sequence of partitions $P=\{P_1,P_2,\dots\}$, with $P_i$ refining $P_{i-1}$, the total information in the parts of $P_n$ is
$$
\tilde S_X^P(n)=\sum_{\chi\in P_n} H(\chi),
$$
and the discrete complexity profile is the marginal gain
$$
C_X^P(n)=\tilde S_X^P(n)-\tilde S_X^P(n-1),\qquad \tilde S_X^P(0)\equiv0.
$$
This profile satisfies nonnegativity and finite support, and it obeys the sum rule
$$
\sum_{n=1}^{\infty} C_X^P(n)=\sum_{x\in X} H(x).
$$
The sum rule is interpreted as a tradeoff between smaller- and larger-scale degrees of freedom: a system cannot be arbitrarily complex at all scales simultaneously [2206.04896].

The key theorem is a multi-scale generalization of Ashby’s law. If $X$ matches $(Y,f)$, then for every nested partition sequence $P$ of $X$, the pulled-back partition sequence $P^f$ on $Y$ satisfies
$$
C_X^P(n)\ge C_Y^{P^f}(n)\quad\forall n.
$$
This is the stated multi-scale law of requisite variety. The same work also gives an additivity theorem for independent subsystems,
$$
C_{A\cup B}^P(n)=C_A^{P|_A}(n)+C_B^{P|_B}(n)
$$
when $I(A;B)=0$, a replicated-blocks theorem,
$$
C_{m*X}(n)=C_X\bigl(\lceil n/m\rceil\bigr),
$$
and a continuum limit obtained from refinements with component size tending to zero [2206.04896].

This formalism differs from the earlier dependency-based complexity profile in its choice of primitive object. It does not begin with signed irreducible dependencies, but with partition-relative information increments. The two approaches nonetheless share a conservation principle and a scale-by-scale interpretation.

## 5. Structural and perceptual variants for patterns, textures, and images

Another major line defines multiscale complexity through coarse-graining and overlap between neighboring renormalized layers. For a pattern $S_0$, one constructs a hierarchy $S_1,S_2,\dots,S_N$ by block-averaging or an RG map. The overlap between successive layers is
$$
O_{k,k+1}=\langle S_{k+1}^{\uparrow},S_k\rangle,
$$
and the per-scale dissimilarity is
$$
C_k=\Bigl|\,O_{k,k+1}-\tfrac12\bigl(O_{k,k}+O_{k+1,k+1}\bigr)\Bigr|.
$$
The total multiscale complexity is
$$
C=\sum_{k=0}^{N-1} C_k.
$$
This measure was reported to vanish for both perfectly ordered and purely random patterns, to peak on visually rich structures, and to detect phase boundaries from a single snapshot, with computational cost $O(L^d\log L)$ [2003.04632].

The same structural idea was adapted to visual stimuli as Multi-Scale Structural Complexity (MSSC). There, a multiresolution hierarchy $\{P_0,P_1,\dots,P_K\}$ is constructed by low-pass filtering, and the per-scale partial complexity is
$$
C_k
=
\Bigl|\langle f_k|f_{k+1}\rangle-\tfrac12[\langle f_k|f_k\rangle+\langle f_{k+1}|f_{k+1}\rangle]\Bigr|
=
\tfrac12\int_D [f_{k+1}(x)-f_k(x)]^2\,dx.
$$
The total MSSC is
$$
C=\sum_{k=0}^{K-1} C_k.
$$
For grayscale images computed via FFT-based filtering, the stated computational complexity is $O(S\cdot N^2\log N)$ [2408.04076].

Applied to the SAVOIAS dataset, “middle-scale” MSSC yielded the following Pearson correlations with subjective complexity scores:

| Category | $\rho$ |
|---|---:|
| Scenes | 0.62 |
| Objects | 0.46 |
| Suprematism | 0.76 |
| Interior design | 0.60 |
| Advertisements | 0.52 |
| Art | 0.36 |
| Infographics | 0.38 |

The same study reported that human judgments rely most strongly on middle spatial bands rather than the finest or coarsest details [2408.04076].

A related image formalism is the multiscale two-dimensional complexity–entropy causality plane. It treats the spatial lags $(\tau_x,\tau_y)$ as the scale, constructs an ordinal-pattern distribution $P^{(S)}$, and then computes the normalized permutation entropy
$$
H_S=\frac{S[P^{(S)}]}{\ln M}
$$
and permutation statistical complexity
$$
C_S=H_S\,Q_J[P^{(S)},P_e].
$$
Varying $(\tau_x,\tau_y)$ produces a trajectory through the $(H_S,C_S)$ plane and was proposed as a multiscale texture descriptor capable of revealing hidden spatial correlations and roughness crossovers [1609.01625].

A common misconception is that randomness is automatically “most complex.” These structural and perceptual formalisms were explicitly motivated by the opposite observation: pure randomness can score highly under entropy or compression criteria while remaining visually less complex than patterns that balance order and disorder [2408.04076].

## 6. Causal and observer-dependent generalizations

Recent work recast multiscale complexity in explicitly causal terms. In one formulation, a macroscale is a surjective coarse-graining $\phi:X\to Y$ of a finite Markov chain, and each scale is assigned a causal measure based on interventionist primitives such as sufficiency, necessity, determinism, and specificity. Along a micro-to-macro path $\pi^{(1)}\to\pi^{(2)}\to\cdots\to\pi^{(K)}$, the unique contribution at step $i$ is
$$
\Delta CI_i=CI_{i+1}-CI_i.
$$
Normalizing the positive gains,
$$
p_i=\frac{\Delta CI_i}{\sum_{j=1}^L \Delta CI_j},
$$
the emergent complexity is defined as the Shannon entropy
$$
EC=-\sum_{i=1}^L p_i\log_2 p_i.
$$
In the worked eight-state Markov-chain example, the microscale causal measure was approximately $0.66$, the macro causal measure approximately $1.00$, and the single-step gain approximately $0.34$ bits [2503.13395].

“Engineering Emergence” extended this path-based picture to the full partition lattice. For a partition $\pi$, the macro TPM $T^\pi$ is assigned determinism and degeneracy, combined into a causal-primitives score
$$
\mathrm{CP}(T^\pi)=\mathrm{determinism}_{T^\pi}-\mathrm{degeneracy}_{T^\pi},
$$
and the non-redundant contribution of a scale is
$$
\Delta\mathrm{CP}(\pi)=\mathrm{CP}(T^\pi)-\max_{\rho<\pi}\mathrm{CP}(T^\rho).
$$
Only partitions with $\Delta\mathrm{CP}(\pi)>0$ belong to the emergent hierarchy. The distribution of gains is then summarized by a path entropy $H_{\rm path}$ and a row entropy $S_{\rm row}$, giving a taxonomy of top-heavy, bottom-heavy, mesoscale-peaked, and scale-free/complex hierarchies. The same work described brute-force enumeration as suitable for $n\lesssim8$ and a branching-greedy sampling method as scaling to $n\sim40$ [2510.02649].

An observer-dependent alternative is Complexity as Advantage (CAA). For an observer class $\mathcal A$, asymptotic average loss is
$$
L(A;X)=\limsup_{|\Lambda|\to\infty}\frac1{|\Lambda|}\sum_{u\in\Lambda}\ell(\hat y_u^A,X_u),
$$
optimal loss is $L^*(X)=\inf_{A\in\mathcal A}L(A;X)$, and regret is $R(A;X)=L(A;X)-L^*(X)$. Complexity is then defined through the variance of regret across observers or the maximal regret gap. Under log-loss and order-$m$ Markov observers,
$$
\Delta L_m=L(A^{(m-1)};X)-L(A^{(m)};X)
=I\bigl(X_t;X_{t-m}\mid X_{t-1}^{\,t-m+1}\bigr),
$$
and
$$
\sum_{m=1}^{\infty}\Delta L_m = E,
$$
where $E$ is excess entropy or predictive information. In this sense, multiscale structure is represented by the profile of predictive gains unlocked by increased observer resources [2511.04590].

These causal and observer-relative approaches do not replace the earlier information-theoretic formalisms. They shift the emphasis from “how much information exists at a scale” to “what a scale uniquely contributes to causal explanation or prediction.”

## 7. Representative applications and characteristic regimes

The complexity-profile formalism has been applied directly to two-dimensional ferromagnetic Ising systems. For a finite $N$-spin system with Hamiltonian
$$
\mathcal H(s)=-J\sum_{\langle i,j\rangle}s_i s_j,
$$
the complexity profile $C(k)$ and scale-specific information $D(k)=C(k)-C(k+1)$ are defined through conditional mutual informations over subsystems, with the scale-weighted sum rule
$$
\sum_{k=1}^N C(k)=N\quad\text{bits}.
$$
The reported behavior separates three regimes: in the disordered phase, $C(1)\approx N$ bits and $C(k)\to0$ for $k\ge2$; near the critical region, a broad spectrum of nonzero $C(k)$ appears, $k=2$ has the largest peak amplitude, and large scales acquire non-monotonic structure, including $C(N-1)$ dipping negative; in the ordered phase, $C(k)\to1$ bit for every $k$ [2507.22060].

The same study examined pairwise complexity densities $c(2)=C(2)/N$ and $d(2)=D(2)/N$ on larger lattices. Both were reported to vanish as $\beta\to0$ and to approach $1$ as $\beta\to\infty$, while each develops a single maximum in the disordered phase below $\beta_c$; the peak height remains bounded in the thermodynamic limit [2507.22060].

In eco-evolutionary dynamics, the formalism was used to argue that patterns of interdependency generate structure at multiple scales and that these structures affect evolutionary trajectories. The thesis presentation connected multiscale structure to spatial host–consumer models, multiplayer games, nonlinear payoffs, ecological stochasticity, and generalized conditionalization rules for replicator dynamics in environments with mesoscale structure [1509.02958].

Across these applications, a stable pattern emerges. Ordered systems tend to concentrate information or causal power at large scales, disordered systems concentrate it at the smallest scales, and systems near transitions or under mixed constraints distribute it across several scales. This suggests a shared diagnostic function for multiscale complexity formalisms: they are designed to locate the scales at which collective organization is genuinely present, rather than to assign a single undifferentiated number to an entire system.

Source: https://www.emergentmind.com/topics/multiscale-complexity-formalism