Papers
Topics
Authors
Recent
Search
2000 character limit reached

Information Emergence Metric (IE): An Overview

Updated 12 July 2026
  • Information Emergence Metric (IE) is a family of info-theoretic models that quantify emergence using measures like normalized Shannon entropy, mutual information differences, and emergent distances.
  • It operationalizes emergence by comparing macro- and micro-level information, revealing differences in contextual integration, complexity, and semantic gain in systems such as large language models.
  • IE also extends to emergent geometries in quantum systems, highlighting the scale-dependence and representation sensitivity inherent in measuring interdependencies and structural unpredictability.

The expression Information Emergence Metric (IE) does not designate a single universally adopted quantity across the arXiv literature. Instead, it refers to several information-theoretic constructions that operationalize “emergence” in different ways: as normalized Shannon information produced by a process (Fernandez et al., 2013), as excess sequence-level mutual information over token-level mutual information inside decoder-only transformers (Chen et al., 2024), and, in a different sense, as an emergent spatial distance induced from quantum mutual information (Leighton-Trudel, 13 Jul 2025). Closely related work uses conditional entropy, binding information, persistent mutual information, metric emergence, and algorithmic-information residuals as emergence-sensitive quantities rather than a single canonical IE definition (Rodríguez-Falcón et al., 12 Oct 2025, Rosas et al., 2018, Carvalho et al., 2022, Abrahão et al., 2021). This plurality suggests that “IE” is best understood as a family of information-theoretic emergence formalisms rather than one standardized metric.

1. Terminological scope and major formulations

Within the supplied literature, at least three distinct usages are central. In Fernández, Maldonado, and Gershenson, emergence is identified directly with Shannon information, so the emergence metric is exactly normalized entropy over a chosen alphabet (Fernandez et al., 2013). In the language-model setting, “Information Emergence (IE)” is defined as a layerwise difference between macro-level and micro-level mutual information, intended to quantify sequence-level contextual information that exceeds isolated token-level information (Chen et al., 2024). In the XXZ spin-chain work, an information-based “emergent distance” is postulated from quantum mutual information as

dE(A,B)=K0I(A;B),d_E(A,B)=\frac{K_0}{\sqrt{I(A;B)}},

with metric validity depending on the scaling law of I(A;B)I(A;B) (Leighton-Trudel, 13 Jul 2025).

Other papers are explicitly adjacent rather than equivalent. The agent-based-model paper proposes Mean Information Gain (MIG), which is mathematically conditional entropy and is used as a proxy for local structural unpredictability in emergent regime classification (Rodríguez-Falcón et al., 12 Oct 2025). The self-organization paper treats binding information

B(Xt)=H(Xt)j=1NH(XtjXtj)B(\boldsymbol{X}_t)=H(\boldsymbol{X}_t)-\sum_{j=1}^N H(X_t^j\mid \boldsymbol{X}_t^{-j})

as the scalar that captures created interdependence, with self-organization defined by the increase of B(Xt)B(\boldsymbol{X}_t) over time (Rosas et al., 2018). Persistent Mutual Information (PMI) is presented as a candidate emergence-sensitive complexity measure because it isolates information that survives across arbitrarily large temporal gaps (Gmeiner, 2012). Metric emergence in continuous dynamics is again something different: not a metric-valued object, but a double-logarithmic complexity exponent derived from quantization in a metric space of probability measures (Carvalho et al., 2022).

A plausible implication is that the topic is unified less by a single formula than by a common program: emergence is rendered measurable by asking how much information is produced, retained, shared, or made geometrically operative.

2. Emergence as Shannon information

The most direct and explicit IE definition in the supplied corpus is the one that sets emergence equal to Shannon information. Fernández, Maldonado, and Gershenson define emergence as “the information a system or process produces” and then state, “Therefore, we can say that emergence is the same as Shannon’s information II” (Fernandez et al., 2013). With probabilities p1,,pnp_1,\dots,p_n, the measure is

E=I=Kipilogpi.E=I=-K\sum_i p_i\log p_i.

The chapter states that from that point onward the logarithm is base 2, and for an alphabet A\mathcal A of size bb, the normalization constant is chosen as

K=1log2b,K=\frac{1}{\log_2 b},

yielding

I(A;B)I(A;B)0

This normalization places the measure in the unit interval:

I(A;B)I(A;B)1

The boundary cases are explicit. If one state has probability I(A;B)I(A;B)2, then I(A;B)I(A;B)3, interpreted as no new information emerging. If the distribution is uniform over the alphabet, then I(A;B)I(A;B)4, interpreted as maximal emergence in the Shannon sense (Fernandez et al., 2013). The measure is computed empirically from symbol frequencies in a discrete time series or, for continuous variables, after discretization into a finite alphabet. The paper stresses that the metric is alphabet-dependent and therefore scale-dependent. The same phenomenon may appear maximally emergent at one scale and non-emergent at another; the example given is the binary string 1010101010, which has I(A;B)I(A;B)5 in base 2 but becomes 22222 with I(A;B)I(A;B)6 in base 4 after grouping pairs (Fernandez et al., 2013).

The 2013 framework places this IE definition inside a larger family of quantities. Self-organization is defined as

I(A;B)I(A;B)7

and complexity as

I(A;B)I(A;B)8

This makes emergence one pole of a polarity rather than a synonym for complexity. In the random Boolean network study, low connectivity I(A;B)I(A;B)9 yields low B(Xt)=H(Xt)j=1NH(XtjXtj)B(\boldsymbol{X}_t)=H(\boldsymbol{X}_t)-\sum_{j=1}^N H(X_t^j\mid \boldsymbol{X}_t^{-j})0, high connectivity yields high B(Xt)=H(Xt)j=1NH(XtjXtj)B(\boldsymbol{X}_t)=H(\boldsymbol{X}_t)-\sum_{j=1}^N H(X_t^j\mid \boldsymbol{X}_t^{-j})1 but low B(Xt)=H(Xt)j=1NH(XtjXtj)B(\boldsymbol{X}_t)=H(\boldsymbol{X}_t)-\sum_{j=1}^N H(X_t^j\mid \boldsymbol{X}_t^{-j})2, and intermediate B(Xt)=H(Xt)j=1NH(XtjXtj)B(\boldsymbol{X}_t)=H(\boldsymbol{X}_t)-\sum_{j=1}^N H(X_t^j\mid \boldsymbol{X}_t^{-j})3 yields balanced B(Xt)=H(Xt)j=1NH(XtjXtj)B(\boldsymbol{X}_t)=H(\boldsymbol{X}_t)-\sum_{j=1}^N H(X_t^j\mid \boldsymbol{X}_t^{-j})4 and B(Xt)=H(Xt)j=1NH(XtjXtj)B(\boldsymbol{X}_t)=H(\boldsymbol{X}_t)-\sum_{j=1}^N H(X_t^j\mid \boldsymbol{X}_t^{-j})5, hence high B(Xt)=H(Xt)j=1NH(XtjXtj)B(\boldsymbol{X}_t)=H(\boldsymbol{X}_t)-\sum_{j=1}^N H(X_t^j\mid \boldsymbol{X}_t^{-j})6 (Fernandez et al., 2013). The same interpretive lesson appears in the Arctic lake case study: high emergence alone indicates continual variation, whereas high complexity requires intermediate emergence rather than maximal emergence. The paper therefore explicitly separates statistical unpredictability from balanced complexity.

3. Information Emergence in LLMs

A distinct use of the term appears in “Quantifying Semantic Emergence in LLMs” (Chen et al., 2024). There, IE is defined internally for decoder-only autoregressive transformers as a blockwise difference between macro-level and micro-level mutual information. The macro variable is a sequence-level representation, typically the last token hidden state in a causal decoder, while the micro variables are token-level representations constructed so that each token depends only on itself. The formal definition is

B(Xt)=H(Xt)j=1NH(XtjXtj)B(\boldsymbol{X}_t)=H(\boldsymbol{X}_t)-\sum_{j=1}^N H(X_t^j\mid \boldsymbol{X}_t^{-j})7

Here B(Xt)=H(Xt)j=1NH(XtjXtj)B(\boldsymbol{X}_t)=H(\boldsymbol{X}_t)-\sum_{j=1}^N H(X_t^j\mid \boldsymbol{X}_t^{-j})8 measures the extent to which a layer transition preserves more information at the sequence level than at the level of isolated tokens (Chen et al., 2024).

The paper motivates this as an entropy-reduction comparison. Mutual information is treated as uncertainty reduction across successive layers, so positive B(Xt)=H(Xt)j=1NH(XtjXtj)B(\boldsymbol{X}_t)=H(\boldsymbol{X}_t)-\sum_{j=1}^N H(X_t^j\mid \boldsymbol{X}_t^{-j})9 means the macro representation undergoes more entropy reduction than the average micro representation. In the authors’ operationalization, this is taken to quantify sequence-level contextual information or “semantic” information emerging from token composition (Chen et al., 2024). The argument is explicitly dynamic: the metric is not B(Xt)B(\boldsymbol{X}_t)0 between macro and micro variables at the same layer, because the paper invokes a supervenience intuition and instead compares macro and micro transitions across layers.

Because hidden states are high-dimensional continuous vectors, the work estimates mutual information using a MINE-style neural critic. The estimator is trained on joint and shuffled pairs of hidden representations across adjacent layers, and the paper reports a 10-layer critic network, batch size B(Xt)B(\boldsymbol{X}_t)1, learning rate decaying from B(Xt)B(\boldsymbol{X}_t)2 to B(Xt)B(\boldsymbol{X}_t)3, and training for B(Xt)B(\boldsymbol{X}_t)4k epochs (Chen et al., 2024). The computational complexity is stated as B(Xt)B(\boldsymbol{X}_t)5, where B(Xt)B(\boldsymbol{X}_t)6 is the cost of estimating one layer-position pair; the appendix reports roughly 40 minutes on one RTX 3090 or 20 minutes on one 4090 for B(Xt)B(\boldsymbol{X}_t)7 (Chen et al., 2024).

The reported empirical patterns are specific. In synthetic in-context-learning settings, IE increases as more shots are added, then often saturates around the sixth or seventh shot (Chen et al., 2024). Tokens within the same shot have approximately the same emergence strength, while the standard deviation of IE rises with shot count. In natural 8-token sentence fragments, IE generally increases with token index, consistent with the idea that later tokens integrate more preceding context. The paper also reports that natural sentences have lower average emergence than synthetic ICL prompts, that larger models tend to show higher emergence, and that prompt form and tokenization matter strongly. The authors emphasize that IE is not a direct measure of semantic correctness, factuality, or general intelligence; it is a measure of context-induced representational emergence (Chen et al., 2024).

4. Emergent geometry from information

A third major line uses information not to quantify uncertainty production but to define geometry itself. In the XXZ spin-chain paper, the proposed emergent distance between subsystems B(Xt)B(\boldsymbol{X}_t)8 and B(Xt)B(\boldsymbol{X}_t)9 is

II0

with the piecewise convention that II1 if II2 and II3 if II4 for II5 (Leighton-Trudel, 13 Jul 2025). Mutual information is computed in the standard way from von Neumann entropies,

II6

where

II7

The paper is explicit that this is a postulate rather than a first-principles derivation. Its analytical core is the Metricity Criterion: if mutual information decays as

II8

then

II9

which is metric exactly when p1,,pnp_1,\dots,p_n0; the Euclidean case corresponds to p1,,pnp_1,\dots,p_n1 (Leighton-Trudel, 13 Jul 2025). In the critical XXZ phase at p1,,pnp_1,\dots,p_n2, the numerics are reported as consistent with power-law mutual-information decay inside this metricity window. In the gapped antiferromagnetic phase at p1,,pnp_1,\dots,p_n3, mutual information decays exponentially,

p1,,pnp_1,\dots,p_n4

which implies

p1,,pnp_1,\dots,p_n5

and the paper argues that this violates the triangle inequality (Leighton-Trudel, 13 Jul 2025).

This construction treats emergence as relational geometry encoded in shared information rather than as entropy production or representational gain. It also introduces a fine-graining constraint: single-site mutual information in the critical chain is expected to scale as p1,,pnp_1,\dots,p_n6, which sits at the metricity boundary, whereas block-block mutual information scales as p1,,pnp_1,\dots,p_n7, giving p1,,pnp_1,\dots,p_n8, which does not satisfy the triangle inequality (Leighton-Trudel, 13 Jul 2025). A plausible implication is that, within this ansatz, emergent geometry is carried by fine-grained entanglement rather than arbitrary coarse-grained correlations.

Several neighboring proposals clarify what an IE metric can mean when the exact label is absent or used differently.

The ABM paper proposes Mean Information Gain

p1,,pnp_1,\dots,p_n9

and, in application, computes it over neighboring occupancy states E=I=Kipilogpi.E=I=-K\sum_i p_i\log p_i.0 on a lattice (Rodríguez-Falcón et al., 12 Oct 2025). MIG is thus conditional entropy, used to quantify how unpredictable a cell’s occupancy remains given a neighbor’s occupancy. Low MIG corresponds to order and strong local predictability; high MIG corresponds to disorder and weak local predictability. The reported regime means are E=I=Kipilogpi.E=I=-K\sum_i p_i\log p_i.1 for convergent, E=I=Kipilogpi.E=I=-K\sum_i p_i\log p_i.2 for periodic, E=I=Kipilogpi.E=I=-K\sum_i p_i\log p_i.3 for complex, and E=I=Kipilogpi.E=I=-K\sum_i p_i\log p_i.4 for chaotic dynamics (Rodríguez-Falcón et al., 12 Oct 2025). The paper therefore uses an information-theoretic proxy for emergence classification, but not a multiscale or causal theory of emergence.

The self-organization framework of Rosas, Mediano, Ugarte, and Jensen treats emergence as the creation of statistical interdependencies. Its principal scalar is binding information,

E=I=Kipilogpi.E=I=-K\sum_i p_i\log p_i.5

which the paper describes as a metric of global structural strength (Rosas et al., 2018). Self-organization is operationally defined by E=I=Kipilogpi.E=I=-K\sum_i p_i\log p_i.6 being an increasing function of time. The paper further decomposes E=I=Kipilogpi.E=I=-K\sum_i p_i\log p_i.7 by extension and sharing mode, distinguishing redundancy-dominated from synergy-dominated structure. Rule 30 in elementary cellular automata is highlighted as having almost zero pairwise mutual information yet large E=I=Kipilogpi.E=I=-K\sum_i p_i\log p_i.8, while rules 60 and 90 realize pure highest-order synergy with

E=I=Kipilogpi.E=I=-K\sum_i p_i\log p_i.9

after two steps (Rosas et al., 2018). This is an emergence-sensitive measure of created dependence rather than of entropy alone.

PMI approaches emergence through long-term informational persistence. It is defined for stationary stochastic processes as

A\mathcal A0

and equivalently as

A\mathcal A1

when the limit exists (Gmeiner, 2012). The paper proves

A\mathcal A2

and

A\mathcal A3

It reports A\mathcal A4 for i.i.d. processes, finite-order Markov processes, and the one-dimensional Ising model, while periodic processes satisfy A\mathcal A5 and the Thue–Morse process yields A\mathcal A6 (Gmeiner, 2012). The argument is that PMI is stricter than excess entropy because it retains only information that survives across arbitrarily large temporal gaps.

Two further notions broaden the scope. Metric emergence in continuous dynamics is defined from quantization of empirical measures, not from Shannon or algorithmic information directly; for a measure A\mathcal A7, the asymptotic exponent is

A\mathcal A8

while topological emergence is the upper metric order of the set of ergodic invariant measures (Carvalho et al., 2022). Algorithmic-information work instead defines observer-dependent emergence as residual conditional algorithmic information after observation and formal theory are accounted for:

A\mathcal A9

and asymptotically observer-independent emergence when that residual eventually exceeds every finite observer-side bound (Abrahão et al., 2021).

6. Limitations, distinctions, and points of controversy

A recurrent limitation is that “information emergence” is not synonymous with “complexity.” In the Shannon-identification framework, high bb0 can correspond to pseudorandomness or disorder, while complexity peaks at intermediate bb1 because bb2 (Fernandez et al., 2013). The ABM paper makes the same point operationally: complex and chaotic regimes both have high MIG, so MIG robustly distinguishes ordered from disordered behavior but is less effective at separating finer structure within those pairs (Rodríguez-Falcón et al., 12 Oct 2025). The PMI paper likewise treats persistent structure as emergence-relevant, but not every persistent pattern need count as emergence in a stronger philosophical sense (Gmeiner, 2012).

A second issue is dependence on representation, scale, or subsystem choice. The Shannon-based emergence metric depends on alphabet size, discretization, and temporal or spatial grouping, and the authors treat this as fundamental rather than incidental (Fernandez et al., 2013). The LLM IE metric depends on fixed sequence length, tokenization, macro/micro variable definitions, and the quality of mutual-information estimation (Chen et al., 2024). The emergent-distance proposal depends crucially on which subsystems are compared: single-site mutual information in the critical XXZ chain supports metricity, while block-block mutual information does not (Leighton-Trudel, 13 Jul 2025). This suggests that any encyclopedia-level definition of IE must distinguish intrinsic claims from representation-relative operationalizations.

A third issue concerns estimation and computability. Shannon-type IE is straightforward once probabilities are estimated, but histogram-based estimation inherits standard sampling caveats (Fernandez et al., 2013). The LLM IE framework relies on MINE-style MI estimation in high-dimensional hidden spaces, which the authors explicitly identify as a bottleneck (Chen et al., 2024). Algorithmic-information emergence is exact only at the level of Kolmogorov complexity and therefore not computable exactly; the paper stresses approximation through compression, coding-theorem methods, block decomposition, and perturbation analysis (Abrahão et al., 2021).

Finally, the literature differs on what emergence is supposed to denote. Some papers equate it with produced information (Fernandez et al., 2013), some with excess sequence-level contextual information (Chen et al., 2024), some with local conditional uncertainty used for regime classification (Rodríguez-Falcón et al., 12 Oct 2025), some with persistent temporal memory (Gmeiner, 2012), some with multivariate interdependence (Rosas et al., 2018), and some with residual algorithmic irreducibility relative to an observer (Abrahão et al., 2021). The XXZ work adds yet another sense: geometry emergent from mutual information (Leighton-Trudel, 13 Jul 2025). The most neutral synthesis is therefore that the Information Emergence Metric is a research area organized around information-theoretic formalizations of emergence, not a single agreed-upon invariant.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Information Emergence Metric (IE).