---
title: 'TensFormer: Hierarchical Occupancy Forecasting'
url: https://www.emergentmind.com/topics/tensformer
type: topic
---

# TensFormer: Hierarchical Occupancy Forecasting

Searching arXiv for the exact and closely related terms to ground the article in current papers.
arXiv search query: "TensFormer OccTENS temporal next-scale prediction"
TensFormer is the generative transformer at the core of OccTENS, a 3D occupancy world model for autonomous driving. It is designed to predict future occupancy scenes and ego motion from historical observations after those observations have been discretized by a scene tokenizer and a motion tokenizer. Its central design choice is to replace flat next-token autoregression with **Temporal Next-Scale Prediction (TENS)**, which factorizes generation into **temporal scene-by-scene prediction** and **spatial scale-by-scale generation**. In the formulation reported for OccTENS, TensFormer is intended to address the inefficiency, temporal degradation, and weak controllability attributed to prior autoregressive occupancy world models [2509.03887].

## 1. Problem setting and representational scope

TensFormer operates in the setting of **occupancy world modeling**, where the target is a sequence of future 3D occupancy scenes rather than ordinary image frames. The motivating difficulty is that occupancy generation must capture fine-grained 3D geometry, semantic occupancy structure, temporal evolution, long-range scene coherence, and ego-motion dependence. The model is therefore embedded in a pipeline with two tokenizers. A **scene tokenizer** converts a 3D occupancy scene into **multi-scale discrete tokens**, and a **motion tokenizer** discretizes ego-motion through relative pose \((x,y,\theta)\). TensFormer then performs sequence generation over these occupancy and motion tokens [2509.03887].

In the reported formulation, a scene \(\mathbf{S}\) is encoded into a BEV latent feature map
\[
\mathbf{F} \in \mathbb{R}^{H \times W \times C},
\]
after which a **multi-scale quantizer** discretizes \(\mathbf{F}\) into
\[
\mathbf{F} = (\mathbf{f}^1,\mathbf{f}^2,\dots,\mathbf{f}^M),
\]
where each \(\mathbf{f}^m\) corresponds to one spatial scale. This is explicitly contrasted with local VQ tokenization that assigns codebook entries only to small local patches; the reported argument is that multi-scale tokenization better preserves global context. Motion is discretized separately and inserted into the same autoregressive sequence as a token \(\mathbf{f}_t^0\), so that each frame is represented by
\[
\{ \mathbf{f}_t^0, \mathbf{f}_t^1,\dots,\mathbf{f}_t^M \}.
\]

The paper does not describe TensFormer as a generic tensorized Transformer in the sense of tensor decomposition or multilinear attention. Its domain specificity is instead architectural and generative: it is a hierarchical Transformer for occupancy forecasting, with explicit handling of time, scale, and motion in a unified sequence model [2509.03887].

## 2. Temporal Next-Scale Prediction

The defining formalism of TensFormer is **Temporal Next-Scale Prediction (TENS)**. Standard autoregressive modeling over a flattened \(n \times n\) feature map is written as
\[
p(x)=\prod_{i=1}^{n\times n} p(x_i \mid x_1,x_2,\ldots,x_{i-1}),
\]
which corresponds to next-token factorization over a single long sequence. OccTENS reformulates this problem so that temporal dependencies are modeled **scene by scene**, while spatial detail is modeled **scale by scale** [2509.03887].

The temporal factorization is given as
\[
p(\mathbf{F}_1,\dots,\mathbf{F}_T)=\prod_{t=1}^{T} p(\mathbf{F}_t \mid \mathbf{F}_0,\dots,\mathbf{F}_{t-1}),
\]
where \(\mathbf{F}_t\) denotes the tokenized representation of frame \(t\). Within a frame, scales are generated autoregressively:
\[
p(\mathbf{F'}_{t-1}, \hat{\mathbf{f}_t^1},\ldots,\hat{\mathbf{f}_t^M}) = \prod_{m=1}^{M} p\!\left( \hat{\mathbf{f}_t^m} \mid \mathbf{F'}_t, \hat{\mathbf{f}_t^1},\ldots,\hat{\mathbf{f}_t^{m-1}} \right).
\]
After motion is integrated, the joint occupancy-motion model becomes
\[
p(\mathbf{F'}_{t-1}, \hat{\mathbf{f}_t^0},\hat{\mathbf{f}_t^1},\ldots,\hat{\mathbf{f}_t^M}) = \prod_{m=0}^{M} p\!\left( \hat{\mathbf{f}_t^m} \mid \mathbf{F'}_t, \hat{\mathbf{f}_t^0},\hat{\mathbf{f}_t^1},\ldots,\hat{\mathbf{f}_t^{m-1}} \right).
\]

The stated purpose of this factorization is to separate **temporal causality** from **intra-frame spatial dependency**. The paper also notes that a naive next-scale adaptation with \(M\) scales for an \(n \times n\) feature map has approximate total token count
\[
n^2\sum_{m=1}^{M}\left(\frac{m}{M}\right)^2 \approx o(Mn^2),
\]
and uses this discussion to motivate a structured decomposition rather than a flat extension of ordinary autoregression. In the reported interpretation, coarse scales capture global scene layout, while finer scales refine local detail; time modeling then governs scene evolution across frames [2509.03887].

## 3. Architectural organization and attention structure

TensFormer has two main generation components: **temporal scene-by-scene prediction** and **spatial scale-by-scale generation**. The scene-by-scene module is further decomposed into **scale-wise temporal causal attention** and **frame-wise spatial attention**. This decomposition is one of the central architectural claims of the model [2509.03887].

For temporal prediction, the model preserves frame causality:
\[
p(\mathbf{F}_1,\dots,\mathbf{F}_T)=\prod_{t=1}^{T} p(\mathbf{F}_t \mid \mathbf{F}_0,\dots,\mathbf{F}_{t-1}).
\]
The paper argues that plain frame-wise causal attention can induce **structure collapse** when generating large-scale occupancy tokens because it mixes **inter-frame temporal causality** with **intra-frame bidirectional spatial dependency**. TensFormer therefore decouples these dependencies. For tokens \(\mathbf{f}_t^m\) at scale \(m\) in frame \(t\), scale-wise temporal causal attention is restricted to
\[
\{\mathbf{F}_1,\mathbf{F}_2,\ldots,\mathbf{F}_{t-1},\mathbf{f}_t^1,\mathbf{f}_t^2,\ldots,\mathbf{f}_t^{m-1}\},
\]
while frame-wise spatial attention uses full attention within the current frame. Spatial refinement then proceeds through a **block-wise causal attention mask** so that each token at scale \(m\) attends only to its prefix:
\[
\{ \mathbf{F}_{t-1}, \hat{\mathbf{f}_t^1},\ldots,\hat{\mathbf{f}_t^{m-1}} \}.
\]

The architectural details reported explicitly are limited. TensFormer contains **three blocks**, with **4 layers each**, hidden dimension **128**, and **4** attention heads. Three 1-D sine-cosine embeddings are added: **position embedding**, **scale embedding**, and **time embedding**. The paper does not provide detailed equations for Q/K/V projections, FFN structure, normalization style, or exact per-scale token layout, and it explicitly notes that some low-level block internals are not specified in the text [2509.03887].

## 4. Pose aggregation and controllable generation

A distinctive aspect of TensFormer is the treatment of ego motion as part of the same token sequence as occupancy. Motion is encoded from relative pose \((x,y,\theta)\), with the \(z\)-axis ignored, and the tokenizer maps it to an embedding
\[
\mathbf{P} = \mathcal{E}(x + y \times V_x + \theta \times V_x \times V_y).
\]
In the resulting sequence, the motion token is treated as the **0-th scale token**:
\[
\mathbf{f}_t^0,
\]
so that each frame is represented as
\[
\{ \mathbf{f}_t^0, \mathbf{f}_t^1, \dots, \mathbf{f}_t^M \}.
\]
This design is termed **holistic pose aggregation** or **multi-modal camera pose aggregation** in the supplied description [2509.03887].

The reported significance of this arrangement is that occupancy generation and ego-motion prediction are not handled in separate branches. Instead, they are modeled through one unified autoregressive sequence. The paper states that this enables two modes of use: conditioning occupancy generation on a specified trajectory, or predicting motion jointly for planning. The relevant joint factorization is
\[
p(\mathbf{F'}_{t-1}, \hat{\mathbf{f}_t^0},\hat{\mathbf{f}_t^1},\ldots,\hat{\mathbf{f}_t^M}) = \prod_{m=0}^{M} p\!\left( \hat{\mathbf{f}_t^m} \mid \mathbf{F'}_t, \hat{\mathbf{f}_t^0},\hat{\mathbf{f}_t^1},\ldots,\hat{\mathbf{f}_t^{m-1}} \right).
\]

Qualitatively, the paper reports that manipulated camera pose inputs yield occupancy generations consistent with turning and lane changing. The mechanism of controllability is therefore sequence-level conditioning through joint motion-occupancy token modeling, rather than an external controller or a separate conditioning network [2509.03887].

## 5. Training setup, objectives, and empirical behavior

OccTENS trains TensFormer in a **two-stage** pipeline: tokenizer training followed by world-model training. For the scene tokenizer, the loss is
\[
\mathcal{L} = \lambda_1 \mathcal{L}_{ce} + \lambda_2 \mathcal{L}_{lovasz} + \lambda_3 \mathcal{L}_{geoscal} + \lambda_4 \mathcal{L}_{semscal},
\]
with
\[
\lambda_1=10.0,\quad \lambda_2=1.0,\quad \lambda_3=0.3,\quad \lambda_4=0.5.
\]
For TensFormer itself, the world model loss is
\[
\mathcal{L} = \beta_1 \mathcal{L}_{occ} + \beta_2 \mathcal{L}_{pose},
\]
with
\[
\beta_1 = 1.0,\qquad \beta_2 = 1.0.
\]
Both \(\mathcal{L}_{occ}\) and \(\mathcal{L}_{pose}\) are cross-entropy losses over discrete tokens [2509.03887].

The reported setup uses **2-second historical context = 4 frames**, predicts **3-second future = 6 frames**, applies occupancy downsampling factor **8**, and uses codebook size **4096**, codebook embedding dimension **128**, and **6 scales** with \([1, 5, 10, 15, 20, 25]\). The world model uses the architectural settings already noted: **3 blocks**, **4 layers each**, hidden dimension **128**, and **4** heads.

The main forecasting results are reported on **nuScenes 4D occupancy forecasting**. Using occupancy as input, the average scores are:
- **OccWorld-O**: mIoU \(17.14\), IoU \(26.63\)
- **OccLLaMA-O**: mIoU \(19.93\), IoU \(29.17\)
- **OccTENS-O**: mIoU \(22.06\), IoU \(31.03\)

Using camera-derived occupancy as input, the average scores are:
- **OccWorld-F**: mIoU \(6.16\), IoU \(18.99\)
- **OccLLaMA-F**: mIoU \(8.66\), IoU \(22.99\)
- **OccTENS-F**: mIoU \(11.79\), IoU \(24.35\)

For planning, the reported averages are:
- **OccWorld**: avg \(L2\) \(1.17\), avg collision \(0.60\%\)
- **OccLLaMA**: avg \(L2\) \(1.14\), avg collision \(0.49\%\)
- **OccTENS**: avg \(L2\) \(1.12\), avg collision \(0.48\%\)

Efficiency is reported through a scale-number ablation. Latencies are:
- **OccWorld**: \(0.35\) s
- **OccTENS, 2 scales**: \(0.21\) s
- **OccTENS, 4 scales**: \(0.34\) s
- **OccTENS, 6 scales**: \(0.56\) s
- **OccTENS, 8 scales**: \(0.93\) s
- **OccSora**: around \(20\) s

The paper uses these results to argue for a controllable quality-efficiency trade-off through the number of scales: **2 scales** is the fastest and already stronger than OccWorld, while larger numbers of scales increase fidelity at increased latency. It also reports qualitatively that **OccWorld exhibits repetition artifacts**, whereas **OccTENS produces more diverse and realistic occupancy scenes** and maintains stronger long-term coherence [2509.03887].

## 6. Position within transformer research and terminological boundaries

Within the supplied sources, the exact term **TensFormer** appears as the internal Transformer of OccTENS [2509.03887]. This should be distinguished from several adjacent but non-identical uses of tensor-oriented or similarly named Transformer work.

**TEAFormer** is a **TEnsor-Augmented Transformer** for multi-dimensional time series forecasting. It preserves matrix- and tensor-variate structure through tensor expansion and Tucker-based compression inside existing time-series Transformer pipelines. The supplied description explicitly states that TEAFormer is related to the query “TensFormer” in the intuitive sense of a **tensor-aware Transformer**, but also states that the paper’s exact term is **TEAFormer**, not TensFormer [2410.20439].

**TensorLens** is not a new Transformer architecture, but a tensor-based formulation for analyzing an existing Transformer as a single input-dependent linear operator encoded by a 4th-order attention-interaction tensor
\[
\mathcal{T} \in \mathbb{R}^{L \times D \times L \times D}.
\]
Its contribution is a global tensor view of attention, FFNs, activations, LayerNorm, residuals, and embeddings, rather than a new generative model for occupancy forecasting [2601.17958].

Accordingly, TensFormer in the strict sense is not a general-purpose tensorized Transformer, nor an analysis framework, nor a Tucker-compressed sequence model. It is a domain-specific hierarchical generator that combines **temporal scene-by-scene prediction**, **spatial scale-by-scale generation**, and **holistic pose aggregation** for occupancy world modeling [2509.03887]. The broader literature nevertheless places it within a family of Transformer research that uses additional structure—scale, tensor organization, or high-order operators—to address weaknesses of flat sequence modeling.

Source: https://www.emergentmind.com/topics/tensformer