---
title: 'HCLSM: Hierarchical Causal Latent State Machines'
url: https://www.emergentmind.com/topics/hclsm
type: topic
---

# HCLSM: Hierarchical Causal Latent State Machines

HCLSM, short for **Hierarchical Causal Latent State Machines**, is a world-model architecture for video-based future-state prediction that is designed around three interconnected principles: **object-centric decomposition**, **hierarchical temporal dynamics**, and **causal structure learning**. It is presented as a response to flat latent representations that entangle multiple scene elements, compress all temporal variation into a single scale, and do not expose interaction structure in a form suitable for reasoning, planning, or counterfactual analysis. In HCLSM, a ViT-based perceptual front end, Slot Attention with spatial broadcast decoding, a three-level temporal engine, and graph-based interaction modeling are combined within a single differentiable system [2603.29090].

## 1. Conceptual motivation and problem formulation

HCLSM is motivated by a specific diagnosis of current world models: many latent-prediction systems can achieve strong predictive performance while representing a scene in a **flat, entangled latent vector**. The architecture identifies three limitations in that regime. First, **object entanglement** prevents clean separation of scene elements such as an acted-on object and a static background. Second, **temporal flattening** obscures the distinction between continuous motion, discrete events such as contact, and long-range goal structure. Third, the absence of **explicit causality** makes it difficult to represent which entity influenced which other entity, or to support robust “what if” analysis [2603.29090].

The paper frames these limitations as a mismatch between representation and physical structure. Its central thesis is that a world model should be structurally aligned with the world it models: **objects are discrete, dynamics are hierarchical, and interactions are causal**. This suggests that next-state prediction alone is an insufficient organizing principle if the objective is not only prediction error minimization but also interpretable latent structure and planner-friendly dynamics.

A related interpretive point in the paper is that lower predictive loss can coexist with poorer latent structure. In HCLSM’s own experiments, a version without spatial broadcast decoding attains lower prediction loss than the two-stage structured version, but the lower-loss model is described as relying on a **distributed, entangled representation**. This is a recurring theme in the design: structural inductive bias is treated as a deliberate constraint rather than as an automatic by-product of predictive training.

## 2. Architectural organization

The system is organized into **five layers**, with experiments focused mainly on the object and dynamics stack rather than on the full continual-memory formulation [2603.29090].

| Layer | Function |
|---|---|
| Perception | ViT encoder over video frames |
| Object decomposition | Slot Attention, spatial broadcast decoding, GNN interaction layer |
| Hierarchical dynamics | Selective SSM, sparse transformer, goal-compression transformer |
| Causal reasoning | Learned adjacency / interaction graph |
| Continual memory | Hopfield + EWC, described conceptually |

The perceptual front end takes a video tensor
$$
(B, T, C, H, W)
$$
and encodes frames into patch embeddings
$$
(B, T, M, d_{\text{model}})
$$
with
$$
M = (H/p)^2,\qquad p=16.
$$
These embeddings are then projected into a world dimension before object decomposition begins.

This layerwise organization is important because HCLSM does not treat perception, object formation, temporal modeling, and interaction reasoning as interchangeable operations inside a monolithic backbone. Instead, each stage is assigned a distinct representational role. A plausible implication is that the architecture is meant to encourage identifiable failure modes: slot collapse, event misdetection, and graph degeneration can each be inspected separately rather than being absorbed into a single opaque latent space.

## 3. Object-centric decomposition

The object-centric component begins with **Slot Attention**, which groups patch tokens into a fixed number of slots $N_{\max}$. Slots are initialized from a learned Gaussian and refined iteratively. The slot competition rule is
$$
\mathbf{A}_{nk} = \frac{\exp(\mathbf{q}_n \cdot \mathbf{k}_k / \sqrt{d})}{\sum_{n'} \exp(\mathbf{q}_{n'} \cdot \mathbf{k}_k / \sqrt{d})}.
$$
This induces competition among slots for explanatory ownership of image tokens [2603.29090].

The model augments standard Slot Attention with an **existence / alive head** $p_{\text{alive}}$ for each slot and a **dynamic birth/death mechanism**, where dormant slots can be instantiated if residual attention energy is high. Empirically, however, the paper reports that this mechanism does not fully solve slot overpopulation: all 32 slots often remain active. That observation is significant because it shows that object-centric factorization is only partial in the current implementation; the architecture promotes decomposition, but it does not yet produce a clean one-object–one-slot correspondence.

To reinforce object structure, HCLSM uses a **Spatial Broadcast Decoder (SBD)** rather than direct pixel reconstruction. Each slot is broadcast over a $14 \times 14$ spatial grid, concatenated with $(x,y)$ coordinates, and decoded by a CNN into feature maps and an alpha mask. The normalized mask competition is
$$
\alpha_{n,p} = \frac{\exp(\hat{\alpha}_{n,p})}{\sum_{n': \text{alive}} \exp(\hat{\alpha}_{n',p})}.
$$
The reconstruction target is not RGB space but a frozen EMA ViT feature, following the DINOSAUR approach:
$$
\mathcal{L}_{\text{SBD}} = \sum_n \sum_p \alpha_{n,p}\, \|\hat{\mathbf{f}}_{n,p} - \mathbf{f}^*_p\|^2.
$$

This design choice is central to the paper’s argument. Reconstructing frozen semantic features rather than pixels is intended to discourage slots from specializing in texture fragments and to encourage them to align with meaningful objects or parts. The paper summarizes the rationale succinctly: **structure must precede prediction**. If temporal objectives dominate too early, the model can minimize loss with distributed latents and never develop explicit spatial ownership.

## 4. Hierarchical temporal dynamics

HCLSM’s temporal engine has **three levels**, each corresponding to a distinct timescale and computational regime [2603.29090].

At **Level 0**, a **selective state space model (SSM)** handles continuous physics. Each object receives its own SSM track with shared parameters across objects. The recurrence is written in Mamba-style form:
$$
\mathbf{h}_t = \overline{\mathbf{A}}_t \odot \mathbf{h}_{t-1} + \overline{\mathbf{B}}_t \odot \mathbf{x}_t,\qquad y_t = \mathbf{C}_t^\top \mathbf{h}_t,
$$
with
$$
\overline{\mathbf{A}}_t = \exp(\Delta_t \mathbf{A}).
$$
A global SSM processes mean-pooled object states and conditions per-object tracks. For numerical stability, $\mathbf{A}_{\log}$ is initialized in $[-0.5, 0]$, and $\Delta_t \mathbf{A}$ is clamped to $[-20, 0]$ to avoid overflow in bf16.

At **Level 1**, a **sparse transformer** handles discrete events such as contact, collision, and sudden state changes. Event detection is based on multi-scale temporal features: frame differences at scales 1, 2, and 4, together with causal dilated convolutions. If an event score exceeds a learned threshold, the corresponding timestep is gathered into a dense event sequence. The computational motivation is explicit: the cost shifts from roughly
$$
\mathcal{O}(T \cdot N^2)
\quad \text{to} \quad
\mathcal{O}(K \cdot N^2),
$$
where $K \ll T$ is the number of detected event steps.

At **Level 2**, a **goal compression transformer** compresses the event stream into abstract summary tokens by means of learned queries and cross-attention. These summaries are fed into a goal-level transformer for long-range reasoning, potentially conditioned on language or goal embeddings.

The hierarchy manager connects these levels by vectorized gather/scatter operations and learned gating. The broader significance of this decomposition is that HCLSM does not treat temporal abstraction as an emergent property of a single backbone. Instead, it explicitly separates smooth motion, eventful transitions, and long-horizon goal structure into distinct computational modules.

## 5. Causal interaction modeling and training curriculum

Causal structure learning in HCLSM begins with a **GNN interaction graph** built over object slots. For each pair of object states, the model forms interaction features
$$
[\mathbf{o}_i;\mathbf{o}_j;\mathbf{o}_i - \mathbf{o}_j;\mathbf{o}_i \odot \mathbf{o}_j].
$$
These are processed by an edge MLP, and messages are aggregated per node. The learned edge weights function as the primary usable interaction signal in the current implementation [2603.29090].

The paper also describes a more explicit causal graph in the form of an adjacency matrix
$$
\mathbf{W} \in \mathbb{R}^{N \times N},
$$
learned by combining Gumbel-softmax binary edge sampling, $\ell_1$ sparsity, and a NOTEARS-style DAG constraint. The stated DAG regularizer is
$$
h(\mathbf{A}) = \mathrm{tr}(e^{\mathbf{A} \odot \mathbf{A}}) - N = 0,
$$
enforced with augmented Lagrangian optimization. However, the paper reports that this explicit causal adjacency did **not** work well empirically: under sparsity regularization, the learned adjacency collapsed to zero. As a result, the GNN edge weights rather than the explicit DAG constitute the operative causal signal. This is an important qualification, because it means that “causal” in the present system is stronger as an architectural aspiration than as a fully validated graph-learning result.

The optimization strategy is a **two-stage training protocol**. In **Stage 1**, which occupies the first 40% of training, only the SBD reconstruction loss and a small diversity regularizer are active; the dynamics loss is computed only for monitoring:
$$
\mathcal{L} \leftarrow 5.0 \cdot \mathcal{L}_{\text{SBD}} + 0.1 \cdot \mathcal{L}_{\text{diversity}}.
$$
In **Stage 2**, covering the remaining 60%, the full prediction objective is enabled while SBD remains as a regularizer:
$$
\mathcal{L} \leftarrow \mathcal{L}_{\text{JEPA}} + \mathcal{L}_{\text{SBD}} + \lambda_{\text{obj}}\mathcal{L}_{\text{obj}} + \lambda_{\text{causal}}\mathcal{L}_{\text{causal}}.
$$

The paper treats this curriculum as one of its major contributions. Its logic is that if all losses are activated from step 0, the JEPA-style prediction term dominates and the model finds shortcuts in entangled latent codes. The two-stage schedule instead forces the system to learn **what things are** before learning **what they do**.

## 6. Empirical evaluation, systems engineering, and limitations

The reported evaluation uses **PushT**, a robotic manipulation benchmark from the **Open X-Embodiment** ecosystem accessed through LeRobot. The dataset contains **206 episodes** and **25,650 frames**. The task involves a robot pushing a **T-shaped block** toward a target with a 2D action space defined by end-effector displacement. Input clips contain **16 frames** at **224 × 224** resolution [2603.29090].

The experimental model is **HCLSM Small**, a **68M-parameter** configuration trained on an **NVIDIA H100 80GB** with batch size 4, learning rate $1.5 \times 10^{-4}$, cosine schedule with 2K warmup, **bf16 mixed precision**, and **50K training steps**, requiring about **6 hours** per run. Stage 1 spans the first 20K steps and Stage 2 the remaining 30K steps. The paper reports **2 successful runs out of 4 launched**, attributing failures to bf16 instability.

The quantitative comparison presented in the paper is as follows.

| Variant | Reported losses | Speed |
|---|---|---|
| HCLSM (no SBD) | Pred. 0.002; Track 0.001; Diversity 0.154; Total 0.100 | 2.3 sps |
| HCLSM (two-stage) | Pred. 0.008; Track 0.016; Diversity 0.132; SBD 0.008; Total 0.262 | 2.9 sps |

The paper interprets this result as a structural tradeoff rather than a simple accuracy ranking. The no-SBD version attains lower prediction loss, but the two-stage version yields **emerging spatial decomposition** and **learned event boundaries**. The event detector typically identifies **2–3 events per 16-frame sequence**, especially around gripper-block contact and other major transitions. PCA visualizations of slot states show trajectories that are smooth most of the time, change direction at event boundaries, and differ across slots.

The implementation is also presented as a substantive contribution. The full system spans **8,478 lines of Python across 51 modules with 171 unit tests**. A **custom Triton kernel** for the selective SSM scan is reported to deliver a large acceleration over a sequential PyTorch implementation. On T4 hardware, the paper gives two examples: **6.22 ms → 0.16 ms**, about **39.3×**, for a tiny configuration, and **69.64 ms → 1.83 ms**, about **38.0×**, for a base configuration. The kernel is said to reduce per-object temporal prediction from the dominant bottleneck to roughly **5% of forward-pass time**. Additional engineering measures include GPU-native tracking via Sinkhorn-Knopp, chunked edge computation when $N > 32$, replacement of `x**2` with `x*x`, activation clamping to $[-50, 50]$, and disabling GradScaler for bf16 on H100.

The paper is equally explicit about limitations. **All 32 slots remain alive**, object decomposition is coarse, and individual objects are split across multiple slots. The explicit DAG-style causal graph collapses to zero under sparsity regularization. Only the **68M** model is trained successfully; larger **262M** and **3B** configurations encounter bf16 NaNs. The evaluation is restricted to a **single dataset with 206 episodes**, and there is **no closed-loop control evaluation**, even though planners are described as implemented.

Taken together, these results position HCLSM less as a finished solution than as a structured research program. Its strongest demonstrated contributions are the explicit integration of slot-based object discovery, multiscale temporal modeling, and graph-mediated interaction reasoning, together with the empirical lesson that object-centric structure may need to be learned before predictive dynamics if entangled latent shortcuts are to be avoided [2603.29090].

Source: https://www.emergentmind.com/topics/hclsm