Papers
Topics
Authors
Recent
Search
2000 character limit reached

Neural Information Squeezer (NIS)

Updated 7 January 2026
  • Neural Information Squeezer (NIS) is a machine learning framework that identifies causal emergence via learned coarse-graining of high-dimensional Markovian systems.
  • It employs a three-module neural architecture—an encoder with invertible transformations, a macro-dynamics learner, and a decoder—to balance information retention with accurate microstate reconstruction.
  • Empirical results in systems like oscillators, Markov chains, and Boolean networks validate NIS's ability to expose multiscale causal structures by maximizing effective information.

Neural Information Squeezer (NIS) is a general machine learning framework for identifying causal emergence through learned coarse-graining of Markovian dynamical systems. The framework employs neural network parameterizations to discover optimal coarse-graining strategies and low-dimensional macro-state dynamics directly from time-series data, maximizing effective information (EI) at the macro level subject to accurate reconstruction of the microdynamics (Zhang et al., 2022). Its architecture explicitly separates information-preserving transformations from information-dropping projections, enabling rigorous analysis of information retention and causal structure across scales.

1. Framework Objective and Problem Setting

NIS addresses the detection and quantification of causal emergence—the phenomenon where a suitable coarse-grained representation of a Markovian system exhibits stronger causal connections than the microscopic description. Given a dynamical system with microstates XtRpX_t \in \mathbb{R}^p and a transition law P(xt+1xt)P(x_{t+1} | x_t), the framework seeks: (a) a differentiable coarse-graining map φ:RpRq\varphi : \mathbb{R}^p \to \mathbb{R}^q, (b) a Markovian macro-dynamics f:RqRqf : \mathbb{R}^q \to \mathbb{R}^q, and (c) a decoder φ\varphi^\dagger that reconstructs microstates from macro ones. The aim is to maximize the effective information in the macro dynamics while maintaining high-fidelity prediction of Xt+1X_{t+1}.

2. Architecture and Information-Preserving Mapping

The NIS architecture comprises three core neural modules:

  1. Encoder (φn\varphi_n): Implements a two-stage transformation φn=χψ\varphi_n = \chi \circ \psi where
    • ψα:RpRp\psi_\alpha: \mathbb{R}^p \to \mathbb{R}^p is an invertible bijection parameterized by stacked RealNVP coupling layers.
    • χpq:RpRq\chi_{p \to q}: \mathbb{R}^p \to \mathbb{R}^q projects onto the first P(xt+1xt)P(x_{t+1} | x_t)0 coordinates, dropping P(xt+1xt)P(x_{t+1} | x_t)1 dimensions.
  2. Macro-dynamics Learner (P(xt+1xt)P(x_{t+1} | x_t)2): Models drift in the macro space as

P(xt+1xt)P(x_{t+1} | x_t)3

so that P(xt+1xt)P(x_{t+1} | x_t)4 is Gaussian.

  1. Decoder (P(xt+1xt)P(x_{t+1} | x_t)5): Inverts P(xt+1xt)P(x_{t+1} | x_t)6, reconstructing microstates as

P(xt+1xt)P(x_{t+1} | x_t)7

where P(xt+1xt)P(x_{t+1} | x_t)8 denotes the concatenation of P(xt+1xt)P(x_{t+1} | x_t)9 and Gaussian noise filling the dropped dimensions.

By explicitly separating information conversion (via bijective φ:RpRq\varphi : \mathbb{R}^p \to \mathbb{R}^q0) from information dropping (via projection φ:RpRq\varphi : \mathbb{R}^p \to \mathbb{R}^q1), NIS enables precise control over the retained information channel width φ:RpRq\varphi : \mathbb{R}^p \to \mathbb{R}^q2 and analytically tracks the information loss.

3. Effective Information and Causal Emergence Calculation

Effective information (EI) of a stochastic map φ:RpRq\varphi : \mathbb{R}^p \to \mathbb{R}^q3 in NIS is defined as the mutual information between a uniform intervention on φ:RpRq\varphi : \mathbb{R}^p \to \mathbb{R}^q4 and the resulting φ:RpRq\varphi : \mathbb{R}^p \to \mathbb{R}^q5: φ:RpRq\varphi : \mathbb{R}^p \to \mathbb{R}^q6 For macro-dynamics φ:RpRq\varphi : \mathbb{R}^p \to \mathbb{R}^q7, an analytical approximation holds: φ:RpRq\varphi : \mathbb{R}^p \to \mathbb{R}^q8 Because φ:RpRq\varphi : \mathbb{R}^p \to \mathbb{R}^q9 diverges, the dimension-averaged effective information f:RqRqf : \mathbb{R}^q \to \mathbb{R}^q0 and the dimension-averaged causal emergence f:RqRqf : \mathbb{R}^q \to \mathbb{R}^q1 provide meaningful comparative metrics.

4. Training Objectives and Information-Bottleneck Regime

NIS training proceeds in two stages:

  • Stage 1 (Reconstruction): For a fixed f:RqRqf : \mathbb{R}^q \to \mathbb{R}^q2, optimize the conditional log-likelihood of observed transitions by maximizing

f:RqRqf : \mathbb{R}^q \to \mathbb{R}^q3

where f:RqRqf : \mathbb{R}^q \to \mathbb{R}^q4 or f:RqRqf : \mathbb{R}^q \to \mathbb{R}^q5, corresponding to Laplace or Gaussian noise respectively. This fits the encoder and macro-dynamics to ensure effective prediction.

  • Stage 2 (Scale Search): Sweep f:RqRqf : \mathbb{R}^q \to \mathbb{R}^q6 over f:RqRqf : \mathbb{R}^q \to \mathbb{R}^q7, retrain, and evaluate f:RqRqf : \mathbb{R}^q \to \mathbb{R}^q8; select f:RqRqf : \mathbb{R}^q \to \mathbb{R}^q9 maximizing emergent EI.

No explicit EI regularization term is required; the architecture's bottleneck structure naturally creates a trade-off between information retention and channel width. Once trained, mutual information φ\varphi^\dagger0 and φ\varphi^\dagger1 converge to φ\varphi^\dagger2 for any φ\varphi^\dagger3; the mapping dynamically allocates useful information per coordinate by scaling φ\varphi^\dagger4.

5. Empirical Demonstrations of Causal Emergence

NIS successfully identifies causal emergence in diverse systems:

System Microstates & Encoding Optimal φ\varphi^\dagger5 Coarse-Graining Effect Causal Emergence
Spring oscillator φ\varphi^\dagger6, φ\varphi^\dagger7 2 Recovers φ\varphi^\dagger8; φ\varphi^\dagger9 matches physics Xt+1X_{t+1}0 peaks at 2
8-state Markov chain 1-hot in Xt+1X_{t+1}1 1 Groups Xt+1X_{t+1}2, Xt+1X_{t+1}3 Macro EI Xt+1X_{t+1}4 micro EI
Boolean network 4 bits, 16 states 1 Clusters 16 microstates into 4 macro groups Matches [Hoel et al.]

In each case, NIS recovers known optimal coarse-grainings and demonstrates positive causal emergence (Xt+1X_{t+1}5) (Zhang et al., 2022).

6. Limitations and Assumptions

Several practical and theoretical limitations arise:

  • RealNVP-based invertible networks for Xt+1X_{t+1}6 are challenging to scale to high-dimensional microstate spaces; stability during training is a concern.
  • Macro transition noise is assumed to be Gaussian (or Laplace); more flexible likelihood models, e.g., normalizing flows on Xt+1X_{t+1}7, are not yet implemented.
  • The coarse-graining map Xt+1X_{t+1}8 is a black-box invertible/projection composite; improving interpretability by imposing sparsity or explicit variable grouping is an open direction.
  • Extensions to continuous-time dynamics (stochastic differential equations) or mappings from trajectories to macro-trajectories are not yet implemented.
  • There is no general closed-form criterion for when causal emergence Xt+1X_{t+1}9 is guaranteed from microdynamics alone; current methodology relies on empirical φn\varphi_n0 computation.

7. Extensions and Future Perspectives

Future work may address scaling the invertible map to higher dimensionalities, incorporating richer probabilistic macro-dynamics, and enforcing interpretable structure in the coarse-grainer. A plausible implication is that, with such enhancements, NIS could systematically uncover multiscale causal structures in complex nonlinear systems from data alone. Establishing theoretical guarantees for causal emergence under broader classes of micro-dynamics remains an open research question (Zhang et al., 2022).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Neural Information Squeezer (NIS).