Papers
Topics
Authors
Recent
Search
2000 character limit reached

Causal Architecture Dynamics Prior to Arrival of Self-replicators in a Model of Catalytic Networks Relevant to Origin-of-Life

Published 30 Jul 2026 in q-bio.PE | (2607.28250v1)

Abstract: Agents that exert causal power in the world are thought to be the product of selection among diverse replicators; what is the causal structure of a medium before replicators appear, and evolution takes hold? We studied information-theoretic dynamics of the popular GARD model that captures several dynamics thought to be important at life's origin and found that causal emergence predicted the initial appearance of self-replication. Moreover, interventions that drove causal emergence up increased the longevity of self-replicators, while interventions that drove causal emergence down decreased the abundance of these self-replicators, suggesting causal emergence as a functional control knob. Thus, progressive increases in integrated causality are detectable in active media before evolutionary dynamics begin to operate, which may have implications for the origin of life across highly diverse scenarios and for understanding the forces driving the rise in causal power observed in the biosphere on Earth.

Summary

  • The paper demonstrates that increased causal emergence predicts the onset of self-replicative dynamics in prebiotic catalytic networks.
  • It applies the GARD model with time-lagged partial information decomposition, confirming notable correlation (73% runs) and significant p-values (<0.001).
  • Intervention experiments establish causal emergence as a 'control knob' that can enhance the persistence and consistency of emerging self-replicators.

Causal Emergence and the Pre-replicator Phase in Catalytic Networks

Introduction

Understanding the transition from prebiotic chemistry to biological self-replication is central to origin-of-life (OoL) studies. Prevailing paradigms ascribe the emergence of agency to natural selection acting on replicators, yet the properties of the underlying pre-replicator medium—particularly its causal architecture—are poorly characterized. This paper analyzes the Graded Autocatalysis Replication Domain (GARD) model, applying information-theoretic measures of causal emergence to probe the dynamics preceding the advent of self-replicators. The findings suggest that increases in integrated causality within catalytic networks both anticipate and modulate the appearance and persistence of self-replicating molecular assemblies.

Methods: GARD Model and Information-Theoretic Analysis

The GARD model simulates the growth and fission cycles of molecular assemblies composed of multiple molecule types, capturing key features hypothesized to underlie prebiotic evolution, notably compositional (as opposed to sequence-based) heredity. Here, the causal emergence metric is computed via a time-lagged partial information decomposition (PID), detecting predictive information generated by the system as a whole but irreducible to its parts. Data preprocessing involves a centered log-ratio transform to address closure constraints in compositional data.

One hundred independent GARD simulations were performed, each tracking assembly composition and classifying system states as either self-replicator or drift, according to established similarity thresholds. The causal emergence measure ("++") was computed for each molecular step, and its temporal properties were compared to network-theoretic and dynamical systems metrics to establish its uniqueness.

Results: Correlation and Predictive Power of Causal Emergence

A key empirical finding is that, while aggregate causal emergence over all runs shows no significant trend, individual runs exhibit pronounced, punctuated spikes in ++ just prior to the emergence of self-replicators. Across runs, 73% show positive correlation between ++ and self-replication, with statistical significance in most cases.

Causal emergence is consistently higher during self-replicating states than in drift states (significant in 57/100 runs; combined p<0.001p < 0.001 via Fisher’s method), and ++ time-series reject white-noise nulls (Ljung-Box p1051p \approx 10^{-51}), demonstrating persistent temporal structure. Notably, machine learning models using early ++ trajectories as inputs significantly outperform baselines in predicting future self-replication, confirming that the information encoded in ++ contains predictive signal regarding the system's transition to self-replicative dynamics.

Interventions and Causality

The pivotal claim of the work is that causal emergence is not only predictive but also mechanistically relevant. Intervention experiments demonstrate that systematically maximizing causal emergence after each assembly fission increases both the persistence (mean 874±233874 \pm 233 steps vs. control 716±198716 \pm 198) and consistency (Pearson ++0 vs. control ++1) of self-replicators, and leads to an accumulating increase in the probability of self-replication over time (++2). Conversely, minimization of ++3 reduces both persistence and overall probability (++4 vs. control ++5) and consistency.

Therefore, ++6 acts as a functional "control knob": modulating the degree of causal emergence within a prebiotic system causally regulates the likelihood and robustness of emergent self-replicators.

Implications and Theoretical Significance

These findings challenge the canonical view that selection acts as the sole engine driving the rise of causal power in biological systems. The detection of increasing integrated causality preceding replication suggests that organizational and information-theoretic ordering principles, distinct from selection, underpin the earliest emergence of agency. This adds quantitative support to metabolism-first and information-based OoL scenarios and highlights the substrate-independence of causal emergence as a theoretically general principle potentially relevant to non-genetic, non-energetic, and even synthetic or artificial chemical systems.

Practically, these insights suggest routes for manipulating the emergence of replicators in engineered or natural systems. Applications may range from steering prebiotic chemistry in vitro, to the design of synthetic protocells, or even controlling dynamical transitions in non-biological complex systems. Further, the paper’s approach unifies dynamical, network, and information-theoretic methodologies, showing limitations of classical graph-based or entropy-based metrics in capturing causal structure relevant for the onset of biological functions.

Future Directions

The results open several lines for future research. First, examining whether analogous causal emergence signatures presage other major evolutionary transitions—e.g., genome compartmentalization or multicellularity—could establish whether this is a general prelude to increases in biological individuality. Second, leveraging ++7 as an order parameter may improve the real-time prediction or even prevention of regime shifts (e.g., in ecosystem collapse or pathological states). Finally, the robust agnosticism of the measure to material implementation suggests utility in AI, nanotechnology, and other domains involving the transition from passive substrates to agentic systems.

Conclusion

This study demonstrates that quantifiable, punctuated increases in causal emergence are detectable in catalytic networks prior to the arrival of self-replicators. These increases both predict and modulate the onset and persistence of self-replicative dynamics. The findings support a view in which integrated causal structure serves as a precursor and enabler of biological agency, independently of genetic heredity or natural selection, and provide operational means to influence the emergence of replicators within complex dynamical systems.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

Overview

This paper asks a simple but deep question: before life had DNA and natural selection, could simple chemical “soups” already show signs of acting together like a team? The authors use computer simulations to study a virtual chemical world and measure how much the whole system acts as one, not just as many separate parts. They show that brief bursts of “working-together power” tend to happen right before simple self-replicating mixtures appear—and that nudging the system to be more integrated can make those replicators last longer.

Key Objectives

The study set out to answer three related questions in everyday terms:

  • What does the “cause-and-effect structure” of a lively chemical soup look like before any self-replicators exist?
  • Do changes in that structure signal when a self-replicator is about to appear?
  • If we push the soup to be more integrated (to “act as one”), can we make self-replication more stable or more likely?

How They Studied It (Methods, in Simple Terms)

They used a well-known origin-of-life simulation called the GARD model. Think of it as a virtual droplet that:

  • Grows by picking up molecules from around it,
  • Reaches a size limit,
  • Splits into two “daughter” droplets,
  • Then keeps repeating this grow-and-split cycle.

Over many cycles, some droplets naturally settle into a repeating mix of molecules. That repeating mix is a simple kind of self-replication: after splitting, the daughters tend to return to the same mix again, like a favorite recipe that keeps being used.

To track “how much the whole acts as a whole,” the authors used an information-theory measure called causal emergence (CE). In plain language:

  • Imagine a sports team. If each player acts alone, the team’s future plays are predictable only from individuals. But if the team coordinates (good passes, shared strategy), knowing the whole team’s pattern tells you more about the next play than looking at players separately. That extra “whole-team” predictability is like CE.
  • Here, the “players” are molecule types, and the “team” is the entire droplet. CE asks: does the whole droplet’s current state tell us more about its future than the parts do on their own?

Key steps they took:

  • They ran 100 independent simulations of the GARD droplet world.
  • At every time step, they measured CE of the current droplet.
  • They also detected when a droplet was in a self-replicating state (a stable, recurring composition).
  • They checked whether CE rose near the times self-replication happened.
  • They trained a simple machine-learning model to see if early CE patterns could predict later self-replication.
  • Finally, they did “interventions”: after each split, they tried adding or removing one molecule type in a way that would specifically raise CE (or, in a separate test, lower CE) and watched how that changed self-replication behavior.

Main Findings and Why They Matter

What they found, in clear points:

  • Spikes matter: Across runs, CE didn’t just steadily rise. Instead, it showed short “spikes,” like quick bursts of teamwork in the droplet.
  • Spikes line up with self-replication: In most simulations, CE was higher during times when the droplet was self-replicating. Many runs showed a positive link between rising CE and being in a self-replicating state.
  • Early warning signal: Using only the early part of the CE curve, a machine-learning model predicted later self-replication better than baselines that used other common measurements (like raw molecule counts or how fast compositions changed). In other words, CE carried special “heads-up” information that the others missed.
  • It’s not just correlation—CE can steer outcomes: When the authors actively tweaked the droplet to maximize CE (by smartly adding or removing a molecule after each split), self-replicators lasted longer and behaved more consistently. Over generations, the chance of being in a self-replicating state also climbed. Doing the opposite—minimizing CE—made self-replication weaker and rarer.
  • Timing beats height: When spikes happened and how far apart they were mattered more than how tall the spikes were. It’s like the schedule of huddles in a team game being more important than how loud the huddle is.

Why this matters:

  • It suggests that even before DNA and natural selection took over, chemical systems could briefly organize themselves in ways that made self-replication more likely.
  • CE captures something genuinely new: it didn’t just mirror standard network statistics, so it provides a fresh “axis” for understanding self-organization.

Implications and Potential Impact

Big-picture takeaways:

  • Origin of life: The jump from “random chemistry” to “self-replication” might be guided by these short bursts of integrated causality. CE spikes could be an early sign that a lifelike process is about to begin.
  • Prediction and control: CE could help scientists predict when simple replicators will appear in different settings—chemical, digital, or even biological. Better yet, gently steering a system to increase CE could help build stable self-replicators when we want them, or avoid them when we don’t.
  • Beyond chemistry: The same idea—“the whole outperforms the parts”—shows up in brains, swarms, ecosystems, and social systems. CE might serve as a general warning light for upcoming “phase changes,” when a system suddenly gains new abilities or organization.

In short, the paper shows that teamwork-like bursts in a chemical soup can both foreshadow and help cause the first steps toward life. That bridges a key gap between inanimate matter and the organized, goal-like behavior we see in living systems.

Knowledge Gaps

Unresolved gaps, limitations, and open questions

Below is a concise, actionable set of gaps that the paper leaves open. Each item highlights what is missing or uncertain and suggests how future work could address it.

Model scope and assumptions

  • Generalizability beyond the original GARD variant is untested; assess whether results hold for updated GARD formulations (e.g., with composome populations, environmental fluxes, kinetic refinements) and other OoL models (RAF theory, RNA-world simulations, digital life systems).
  • Parameter sensitivity is not explored; perform systematic sweeps over Ng, nmin/nmax, maxsteps, fission rules, and the lognormal rate parameters (mean, variance) to map regimes where CE–replication coupling holds or fails.
  • Environmental realism is limited (homogeneous, time-invariant “soup”); test robustness under spatial structure, compartment growth/merger, resource limitation, wet–dry or thermal cycles, and energy/food-source constraints.
  • Population-level dynamics and selection are absent (single lineage tracked); extend to interacting populations of assemblies to test whether CE predicts competitive success, coexistence, and selection-driven transitions.
  • Fission modeled as symmetric binomial split; evaluate alternative, chemically plausible division rules (asymmetric splits, composition-dependent scission) and their impact on CE and replicator persistence.
  • Energetics and thermodynamic consistency are not addressed; quantify links between CE, free-energy dissipation, and non-equilibrium driving to evaluate physical plausibility and costs.

Causal emergence (CE) metric and estimation

  • The CE computation relies on a minimum-information bipartition (MIB) reduction to two components; quantify how results change with different partitions, multiway partitions, or searches (exact vs heuristic) and report computational scaling with Ng.
  • PID choice/implementation is under-specified; compare alternative PID/synergy formalisms (e.g., I_dep, BROJA, O-information, Integrated Information Decomposition variants) and continuous vs discretized estimators to test metric dependence.
  • Estimator bias/variance and finite-sample effects are not quantified; report confidence intervals, null-model baselines (e.g., time-shuffled, composition-permuted, rate-matrix-permuted controls), and power analyses for CE trajectories and spikes.
  • The centered log-ratio (clr) transform plus component removal may distort dependencies; benchmark against isometric log-ratio (ilr) bases and evaluate invariance of CE findings to compositional-data transform choices.
  • Interpretation of negative CE values and the absolute scale is unclear; specify normalization, units, and the theoretical range to aid comparability across runs and models.

Self-replicator definition and detection

  • The detection criterion (Euclidean similarity to the most recurring composition and a threshold) is not justified or stress-tested; perform sensitivity analyses over thresholds, distance metrics (cosine, Aitchison), and clustering algorithms.
  • False positive/negative rates for replicator detection are not reported; quantify detection accuracy with synthetic ground-truth tests (known attractors) and report robustness to noise and short transients.
  • It is unclear whether CE spikes also precede other structural events (non-replicator attractors, composition bottlenecks); analyze event-specificity by labeling and comparing CE dynamics around multiple event types.

Predictive claims and statistics

  • Predictive modeling uses an MLP with limited baselines; compare against time-series models (logistic regression with lags, survival models, HMMs, random forests, sequence models) and report precision/recall, AUC, calibration, and lead-time distributions.
  • Generalization is only tested across runs with identical settings; evaluate out-of-distribution performance across parameter regimes, catalytic matrices, and GARD variants.
  • Lead–lag structure is not directly quantified; compute lagged cross-correlations, Granger causality, and information-theoretic transfer measures to estimate predictive horizons and causal directionality.
  • Multiple-comparisons handling is unclear across many per-run tests; report correction procedures, effect sizes, and run-level versus aggregated inferences that respect between-run dependence structures.
  • Fisher’s method assumes independence; justify or replace with meta-analytic approaches that model between-run heterogeneity.

Interventions and causal interpretation

  • Interventions (add/delete one molecule post-fission chosen to maximize CE) may not be chemically realizable; explore physically plausible intervention classes (e.g., controlled influx rates, catalytic activity modulation, environmental gradients) and compare causal impacts.
  • Only one intervention timing/strength is tested; vary frequency, magnitude, and timing (pre-growth vs post-fission) to map dose–response curves for CE control and replicator outcomes.
  • Mechanistic pathways by which CE increases replicator persistence are not unpacked; analyze how CE-changing interventions alter network motifs, effective autocatalytic closure, and basin-of-attraction structure.
  • Necessity/sufficiency of CE spikes is unproven; quantify P(replicator | CE spike) and P(CE spike | replicator), false alarm/miss rates, and conditions producing spikes without replication (and vice versa).
  • Potential confounds from state perturbations are unaddressed; include sham interventions that preserve composition norms or random molecule changes matched in magnitude to isolate CE-specific effects.

Mechanistic and theoretical integration

  • The origin of CE spikes is not explained; identify micro-to-macro transitions (e.g., percolation of catalytic pathways, transient closure) that produce spikes via motif and RAF-set analyses.
  • Relationship between CE and known dynamical indicators (Lyapunov exponents, entropy rates, correlation dimension) is reported as “non-significant” but not power-assessed; perform controllability/observability analyses to reconcile why CE captures unique structure.
  • Link to thermodynamics and constraints is absent; test whether CE correlates with dissipation, entropy production, or gradients, and whether trade-offs exist between CE and energetic cost.
  • Multi-scale structure is not explored; examine whether coarse-grained macrovariables yield higher predictive power (macro-beats-micro) and whether optimal coarse-grainings align with emergent replicator states.

Validation and external relevance

  • No laboratory validation is presented; test CE measurement and CE-targeted interventions in experimental protocell or lipid-catalytic systems (e.g., vesicles with compositional heredity) to assess empirical feasibility.
  • Cross-model replication is missing; reproduce the CE–replication link in independent OoL frameworks (RAF, stochastic autocatalytic sets, cellular automata, Avida) to evaluate substrate-agnostic claims.
  • Long-term dynamics are truncated (100 generations); extend to longer horizons to assess stability, recurrence of spikes, and late-emerging parasites/error-threshold effects, and whether CE anticipates these transitions.
  • Code and exact algorithms (MIB search strategy, PID implementation details, spike detection criteria) are not yet available; open these resources with unit tests and deterministic seeds to enable full reproducibility.
  • Practical control recipes are undeveloped; translate CE modulation into implementable levers (environmental cycling schedules, feed compositions, catalysts) and quantify intervention cost–benefit for accelerating or suppressing replicator emergence.

Practical Applications

Immediate Applications

Below are near-term, actionable uses that can be piloted or deployed with current data, software, and lab capabilities.

  • Causal-emergence early-warning analytics for complex systems
    • Sector: software, cybersecurity, IT operations (AIOps), industrial IoT
    • What: Monitor multivariate time-series from networks or services and flag CE spikes as precursors of self-reinforcing, hard-to-reverse events (e.g., worm/replicator outbreaks, cascading failures).
    • Tools/products/workflows:
    • CE-Monitor SDK implementing time-lagged PID with minimum-information bipartition (MIB) on streaming telemetry; connectors for SIEM/observability stacks.
    • Alerting policies that treat CE spikes as “pre-incident” signals; MLP classifier trained on CE trajectories to forecast outbreak windows.
    • Assumptions/dependencies: Requires sufficiently rich, synchronized multivariate telemetry; discretization/binning choices and MIB approximations affect sensitivity; false positives possible during benign coordination bursts.
  • In silico steering of digital replicators and open-ended simulations
    • Sector: research software, evolutionary computation, artificial life
    • What: Use CE as a control knob in platforms like Avida/Tierra or agent-based models to promote or suppress emergent self-replication; validate design hypotheses before wet-lab work.
    • Tools/products/workflows: CE-guided intervention loop that adds/removes agents/interactions to maximize/minimize CE after each “generation,” mirroring the paper’s intervention protocol.
    • Assumptions/dependencies: Transferability from GARD-like dynamics depends on the simulator’s causal structure; requires access to internal state variables.
  • OoL and systems-chemistry data analysis
    • Sector: academia (origin-of-life, protobiology), chemical network theory
    • What: Apply CE to existing simulation outputs and time-series from protocell/protometabolic experiments to identify predictive spikes preceding compositional heredity.
    • Tools/products/workflows:
    • Open-source pipelines for compositional data (centered log-ratio transforms), CE estimation, and spike detection; model comparison against entropy, Lyapunov, and network centralities.
    • MLP-based predictors trained on early CE segments to forecast first-appearance windows of self-replicators.
    • Assumptions/dependencies: Requires longitudinal, multivariate composition measurements; pre-processing for compositional closure is critical.
  • Swarm and multi-agent system supervision
    • Sector: robotics, distributed AI, autonomous systems
    • What: Monitor CE of swarm states to detect imminent lock-in to self-reinforcing behaviors (beneficial coordination vs. pathological crowding).
    • Tools/products/workflows: Real-time CE telemetry from agent state vectors; autopilot policy that nudges tasking/communication to increase (for robust coordination) or decrease (to avoid self-trapping) CE.
    • Assumptions/dependencies: Access to internal state or high-fidelity estimators; low-latency CE computation.
  • Neuroscience and clinical monitoring add-on metric
    • Sector: healthcare (neuro), clinical research
    • What: Add CE to EEG/fMRI analytics as a complementary measure of emergent integration (agency-related dynamics), leveraging prior clinical use-cases of information-integration metrics.
    • Tools/products/workflows: CE module in existing neuroinformatics pipelines; validation against task onsets, anesthesia depth, or recovery trajectories.
    • Assumptions/dependencies: Data quality, stationarity windows, and preprocessing choices affect reliability; interpretational care needed (correlation vs causation).
  • Teaching and outreach on emergence
    • Sector: education
    • What: Classroom visualizations showing how CE spikes precede emergent phases in simple agent/chemistry simulators.
    • Tools/products/workflows: Interactive notebooks and dashboards with CE overlays on state-space trajectories.
    • Assumptions/dependencies: Pedagogical framing to avoid overgeneralization beyond modeled systems.

Long-Term Applications

These opportunities are plausible extensions but will need further research, validation in real substrates, scaling, and/or regulatory frameworks.

  • Closed-loop steering of prebiotic-like chemical reactors
    • Sector: systems chemistry, synthetic protocells
    • What: Microfluidic reactors that measure composition in situ and apply CE-maximizing feedstock tweaks post-fission to increase persistence/consistency of compositional self-replicators.
    • Tools/products/workflows:
    • CE-Lab: mass spec or spectroscopic proxies feeding a controller that selects “add/remove” reagent actions that maximize CE, per the paper’s intervention design.
    • Digital twins that learn intervention policies in silico before lab execution.
    • Assumptions/dependencies: Real-time compositional sensing, robust mapping from “molecule type” abstractions to actual chemistries, and actuation latencies compatible with system dynamics.
  • Programmable self-assembly and nanomanufacturing
    • Sector: nanotechnology, materials
    • What: Use CE as an online objective to guide self-assembling colloids, DNA origami, or peptide networks toward robust, self-maintaining structures; suppress unintended self-replication.
    • Tools/products/workflows: CE-driven design-of-experiments; reinforcement learning controllers that modulate environment (pH, salt, feed ratios) to tune CE.
    • Assumptions/dependencies: Proxy observables for “molecular types” and reliable, reversible interventions at nano/micro scales.
  • Managing pathological attractors in biomedicine
    • Sector: healthcare (oncology, fibrosis, chronic inflammation)
    • What: Identify high-CE, self-maintaining physiological states (e.g., tumor microenvironment niches) and design interventions that reduce CE to destabilize pathologic attractors or increase CE to stabilize healthy ones.
    • Tools/products/workflows: Multi-omics time-series or spatial transcriptomics informing CE estimation; in silico screening of CE-lowering perturbations (e.g., cytokine or ECM-targeting combos) before animal/clinical studies.
    • Assumptions/dependencies: Dense longitudinal data from living tissues; causal interpretability in heterogeneous, noisy biology; safety of CE-targeted perturbations.
  • Information-network and social contagion risk governance
    • Sector: policy, platform governance, public health communication
    • What: Treat CE spikes in information flow graphs as early indicators of self-replicating narratives/botnets; apply structural/content interventions to reduce CE and prevent runaway propagation.
    • Tools/products/workflows: CE-Guard integrated into trust-and-safety platforms; playbooks for throttling amplification or altering graph connectivity when CE thresholds are crossed.
    • Assumptions/dependencies: Access to platform telemetry; ethical constraints; avoiding suppression of legitimate coordination.
  • Financial stability early warnings
    • Sector: finance
    • What: Use CE to detect self-reinforcing market dynamics (feedback loops across assets/funds) ahead of regime shifts; trigger graduated circuit-breakers or liquidity backstops.
    • Tools/products/workflows: CE computed on multivariate order-book and cross-asset flow data; governance dashboards with CE trendlines and automated guardrails.
    • Assumptions/dependencies: High-frequency, multi-venue data access; risk of false positives during orderly coordination; regulatory acceptance.
  • Ecosystem and infrastructure resilience
    • Sector: ecology, energy grids, transportation
    • What: CE as a precursor signal of regime shifts (eutrophication, blackout cascades) enabling preemptive load-shedding or restoration actions; conversely, CE-maximizing actions to stabilize desired coordinated operation.
    • Tools/products/workflows: Sensor network fusion into CE estimators; model-predictive control that includes CE targets.
    • Assumptions/dependencies: Adequate sensor coverage and calibration; domain-specific mapping from interventions to CE changes.
  • AI safety and multi-agent alignment
    • Sector: AI research, autonomous systems
    • What: Monitor CE across agent populations or internal modules to detect onset of self-replicating or self-preserving subdynamics; throttle or restructure training when CE crosses bounds.
    • Tools/products/workflows: CE-constraints in open-ended or population-based training; dashboards for CE spikes during exploration.
    • Assumptions/dependencies: Observability of internal states; establishing thresholds that distinguish useful coordination from risky self-replication.
  • Engineered microbial consortia and bioprocess control
    • Sector: biotech, biomanufacturing
    • What: CE-guided tuning of community composition and feed strategies to lock processes into productive, self-maintaining states, or unlock stuck failure modes.
    • Tools/products/workflows: Online metabolomics/flow cytometry feeding CE controllers; CE-aware media/feed optimization.
    • Assumptions/dependencies: Real-time multi-omics; robust mapping from interventions to CE in living consortia.
  • Personal and public health early warnings
    • Sector: daily life, digital health, public health
    • What: Future wearables and home sensors compute CE across multimodal signals (sleep, HRV, glucose, mood logs) to anticipate transitions into hard-to-exit states (e.g., migraine, relapse, burnout) and suggest CE-lowering/raising routines.
    • Tools/products/workflows: On-device CE estimation; nudging interventions (light, activity, nutrition) tailored to push CE in desired directions.
    • Assumptions/dependencies: High-quality, multi-sensor data; individualized model calibration; privacy and clinical validation.

Cross-cutting assumptions and dependencies

  • Generalization: Results were shown in GARD (a compositional, pre-genetic model). Translating to real chemistries, organisms, or socio-technical systems requires empirical validation.
  • Data requirements: CE estimation needs multivariate, time-resolved data; compositional datasets require centered log-ratio transforms or equivalent preprocessing.
  • Computation: Time-lagged PID with MIB can be computationally heavy; scalable approximations and robust discretization are needed for real-time use.
  • Causality vs. detection: CE spikes were predictive and manipulable in silico, but causal levers and side effects in real systems must be established experimentally.
  • Ethics and governance: Using CE to suppress or promote “replicators” (biological or informational) carries safety, fairness, and freedom-of-expression considerations.

Glossary

  • Abiotic-to-biotic transition: The shift from non-living chemical dynamics to living systems capable of processing information. "viewing the abiotic-to-biotic transition as the onset of systems capable of storing, transmitting, and acting upon information."
  • Accretion: Growth by gradual accumulation of external material. "GARD simulates assemblies that grow by accretion of environmental molecules."
  • Attractor: A stable set or state that dynamical trajectories tend to converge to. "like attractors in dynamical systems [84]."
  • Autocatalytic networks: Chemical networks in which molecules catalyze each other’s formation, producing self-sustaining sets. "thermodynamically driven chemical systems [33], autocatalytic networks [34-36], and growth-first scenarios [37]"
  • Betweenness centrality: A network metric quantifying how often a node lies on shortest paths between other nodes. "(number of nodes and edges, in-degree, out-degree, betweenness centrality, PageRank, HITS scores)"
  • Binomial distribution: The probability distribution of the number of successes in a fixed number of independent Bernoulli trials. "sampled from a binomial distribution with parameter 0.5."
  • Causal cut: A partition between system components that minimizes causal/information flow across it. "and hence have the weakest causal cut."
  • Causal emergence: An increase in effective causal power at the system (macro) level beyond what parts predict independently. "We found that causal emergence predicted the initial appearance of self-replication."
  • Causal information theory: An information-theoretic framework for quantifying causation and agency in dynamical systems. "Recent work in causal information theory has significantly advanced our ability to quantify when and how much a given system is an integrated causal agent [12-22]."
  • Causal integration: The extent to which components act as a unified whole to influence future states. "they will have a higher causal integration because of their ability to communicate through pheromones,"
  • Centered log-ratio transform: A transformation for compositional data that removes closure effects by referencing the geometric mean. "we applied the centered log-ratio transform:"
  • Circular causation: Feedback causation where causes and effects influence each other cyclically. "While debates about downward and circular causation continue [1-9],"
  • Co-information: A multivariate information measure capturing shared information among variables (can indicate redundancy or synergy). "Other measures of integrated information exist, such as total correlation and co- information."
  • Compositional heredity: Inheritance based on the composition of molecular types rather than genetic sequences. "They presented compositional heredity without genetic material, showing how life-like reproduction can emerge from purely chemical systems [85]."
  • Compositional inheritance: The transmission of a system’s component proportions across generations instead of sequence-based encoding. "including the emergence of mutually catalytic molecular assemblies and compositional (rather than sequence-based) inheritance."
  • Composome: A stable, self-reproducing molecular composition (assembly) in the GARD framework. "the composomes, shown as colored circles, with one color per molecule type"
  • Correlation dimension: A fractal dimension estimating the complexity of an attractor from a time series. "(sample entropy, correlation dimension, Lyapunov exponents, DFA, generalized Hurst exponent)"
  • Detrended Fluctuation Analysis (DFA): A method to quantify long-range temporal correlations in noisy, nonstationary time series. "sample entropy, correlation dimension, Lyapunov exponents, DFA, generalized Hurst exponent"
  • Downward causation: Higher-level structures or properties influencing the behavior of lower-level components. "While debates about downward and circular causation continue [1-9],"
  • Euclidean space: Standard geometric space used here to measure similarity of compositions. "highly similar (in Euclidean space) from one generation to the next [65]."
  • Fisher's method: A technique for meta-analysis that combines p-values from independent tests. "Using Fisher's method (p<0.001), the combined evidence from the 100 independent runs significantly supported that causal emergence was higher during steps when the assemblies were in self-replication,"
  • Fission: The splitting of one assembly into two daughter assemblies. "Assemblies underwent fission when they reached a critical size,"
  • Generalized Hurst exponent: A scaling exponent characterizing long-range dependence and multifractality in time series. "sample entropy, correlation dimension, Lyapunov exponents, DFA, generalized Hurst exponent"
  • Graded Autocatalysis Replication Domain (GARD): A model of mutually catalytic molecular assemblies exhibiting compositional reproduction. "We chose the popular Graded Autocatalysis Replication Domain (GARD) model [54-64]"
  • HITS scores: Authority and hub scores from the Hyperlink-Induced Topic Search algorithm used to rank nodes in a network. "betweenness centrality, PageRank, HITS scores"
  • Homeostatic growth: Growth that preserves a stable internal composition despite turnover. "self-replicators are specific molecule compositions that exhibit homeostatic growth (attractors)"
  • Integrated causality: The joint causal influence generated by a system considered as a whole. "progressive increases in integrated causality are detectable in active media before evolutionary dynamics begin to operate,"
  • Integrated information: A measure of how much information is generated by the whole system beyond its parts. "Other measures of integrated information exist, such as total correlation and co- information."
  • Ljung-Box test: A statistical test for detecting autocorrelation in time series residuals. "the null hypothesis of temporal independence was rejected in 86/100 runs (median Ljung-Box p=2.07x10-51),"
  • Lognormal distribution: A probability distribution of a random variable whose logarithm is normally distributed. "The rates of the catalytic matrix ß were sampled from a lognormal distribution with mean A and standard deviation o"
  • Lyapunov exponents: Quantities measuring the rates of separation of infinitesimally close trajectories (chaos indicator). "sample entropy, correlation dimension, Lyapunov exponents, DFA, generalized Hurst exponent"
  • Mann–Whitney test: A nonparametric test for assessing differences between two independent samples. "Indeed, ør was significantly higher (using Mann-Whitney, p<0.001)"
  • Minimum-information bipartition: A split of system variables into two parts that share the least information, used to find a weakest link for analysis. "we adopted the minimum-information bipartition [90] to reduce the system to a computationally tractable two-component system."
  • Multilayer perceptron (MLP): A feedforward neural network model with one or more hidden layers. "we fitted an MLP, using 80% of the runs for training and 20% for testing"
  • PageRank: A network centrality algorithm ranking nodes by the importance conferred by incoming links. "betweenness centrality, PageRank, HITS scores"
  • Partial Information Decomposition (PID): A framework that decomposes information into unique, redundant, and synergistic components. "the Partial Information Decomposition [88] and its recent extension to time series like ours, the PID decomposition [82, 89]."
  • Pearson correlation coefficient: A measure of linear association between two variables. "measured as the Pearson's p between consecutive steps;"
  • Poisson process: A stochastic process modeling counts of random events with a constant average rate. "a sequence of stochastic (Poisson) updates"
  • Prebiotic soup: The hypothesized mixture of chemicals on early Earth from which life emerged. "To capture the conditions of the prebiotic soup, the GARD model is essentially a stochastic model of events before natural selection takes place."
  • Sample entropy: A statistic quantifying the complexity or irregularity of a time series. "dynamical properties (sample entropy, correlation dimension, Lyapunov exponents, DFA, generalized Hurst exponent)"
  • Self-replicator: A system or composition that reproduces itself over time, maintaining its defining structure. "Self-replicators emerged spontaneously (as a product of specific GARD dynamics and parameters) as recurring compositions that are "inherited" across generations [64, 65]."
  • Shannon mutual information: A measure of shared information between variables. "is the time-lagged multivariate Shannon mutual information."
  • Simplex: The constrained space of compositions whose components sum to one. "the data live on a simplex, creating spurious correlations."
  • Spearman correlation: A rank-based measure of monotonic association between variables. "we tested this correlation using the Spearman correlation test (Figure 1)."
  • Total correlation: A multivariate measure of overall statistical dependence among variables. "Other measures of integrated information exist, such as total correlation and co- information."
  • White noise: A time series with no serial correlation; purely random fluctuations. "Most ør trajectories rejected the null hypothesis of white noise,"

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 3 tweets with 432 likes about this paper.