---
title: Temporal Next-Scale Prediction (TENS)
url: https://www.emergentmind.com/topics/temporal-next-scale-prediction-tens
type: topic
---

# Temporal Next-Scale Prediction (TENS)

Searching arXiv for the cited TENS-related papers and neighboring work.
Temporal Next-Scale Prediction (TENS) denotes a family of predictive formulations in which a model learns structured forecasts across multiple temporal horizons or temporal resolutions rather than only a single immediate next step. In its earliest reinforcement-learning formulation, the idea appears as multi-timescale nexting: each measurable signal is treated as a pseudo-reward, and a bank of value functions predicts exponentially discounted future signals at distinct short horizons [1112.1133]. In later generative formulations, TENS is used more explicitly to describe multi-scale temporal objectives such as predicting next\(^1\), next\(^2\), and next\(^3\) video chunks with a causal chain [2606.11187], or generating sequences coarse-to-fine across temporal scales in motion and occupancy models [2604.03799; 2509.03887]. Across these usages, the unifying principle is that anticipatory structure is distributed over a set of “next-scales,” with coupling across horizons or resolutions used to improve temporal coherence, controllability, or sample efficiency.

## 1. Origins in multi-timescale nexting

The foundational formulation of TENS is implicit in multi-timescale nexting on a reinforcement-learning robot [1112.1133]. “Nexting” denotes the continual prediction of what will happen next in an immediate, local, and personal sense. In that setting, each measurable signal \(x_t^i\) is treated as a pseudo-reward \(r_t^i\), and the learner estimates a separate value function for each signal and each timescale. TENS, in this sense, is realized by a bank of temporal-difference predictors that share a common feature representation while differing in target signals and discount factors.

For each signal, the prediction target is the exponentially discounted return
\[
G_t^i \equiv \sum_{k=0}^{\infty} (\gamma^i)^k \, r_{t+1+k}^i,
\]
with value function
\[
v^i_\gamma(s_t) = \mathbb{E}\!\left[\sum_{k=0}^{\infty} \gamma^k \, r_{t+1+k}^i \;\middle|\; s_t\right].
\]
Multiple temporal scales are obtained by assigning distinct \(\gamma\) values per prediction. If data are sampled every \(\Delta t\) seconds, the desired continuous-time horizon \(\tau\) is approximated by
\[
\gamma \approx e^{-\Delta t / \tau},
\]
and the effective discrete-time horizon by
\[
H \approx \frac{1}{1-\gamma}.
\]

In the Critterbot experiments, \(\Delta t \approx 0.1\,\mathrm{s}\), and four timescales were used: \(\gamma \in \{0, 0.8, 0.95, 0.9875\}\), corresponding approximately to horizons of \(0.1\) s, \(0.5\) s, \(2\) s, and \(8\) s [1112.1133]. The reported system learned 2160 predictions online, updated at better than \(10\) Hz, with most dramatic gains within approximately \(30\) minutes. Learned predictions for an ambient light sensor at \(2\) s and \(8\) s rose and fell in anticipation of saturation near a lamp, matching empirical returns and the offline-optimal least-squares solution closely. This establishes the canonical interpretation of TENS as short-horizon anticipatory prediction over a spectrum of timescales.

A broader implication suggested by later work is that multi-timescale nexting supplies a general template: select temporal scales, define a prediction target at each scale, and couple the resulting predictors through a shared representation rather than through a single monolithic output.

## 2. Formalizations of temporal next-scale prediction

Across the cited literature, TENS takes two mathematically distinct but conceptually related forms.

The first is the discounted-return formulation of multi-timescale nexting, where different \(\gamma\) values encode different effective horizons [1112.1133]. This formulation preserves a conventional value-function semantics: each predictor estimates an expectation of future signal accumulation over a chosen timescale.

The second is an explicit multi-depth or multi-resolution prediction objective, in which the model predicts several future chunks, frames, or scales jointly. In Next Forcing, TENS is defined as jointly predicting
\[
\{C_{i+1}, C_{i+2}, C_{i+3}\}
=
\{next^1(C_i), next^2(C_i), next^3(C_i)\},
\]
with causal supervision at each depth and causal chaining across depths [2606.11187]. In VideoAR, the factorization is over both time and intra-frame scale:
\[
p(R_{1:T}^{1:K} \mid \Psi) = \prod_{t=1}^{T} \prod_{k=1}^{K} p(R_t^k \mid R_{1:t-1}^{1:K}, R_t^{1:k-1}, \Psi),
\]
so temporal next-scale prediction couples next-frame prediction with next-scale prediction inside each frame [2601.05966]. In MoScale, the factorization becomes
\[
p(X^{(1)},\dots,X^{(K)} \mid c) = \prod_{s=1}^{K} p(X^{(s)} \mid X^{(<s)}, c),
\]
where each scale spans the complete temporal horizon at a different resolution [2604.03799]. In ScaleMoGen, the same pattern appears as
\[
p(M^{(0)}, \ldots, M^{(V)} \mid c) = \prod_{v=0}^{V} p(M^{(v)} \mid M^{(<v)}, c),
\]
with optional bitwise decomposition inside each scale [2605.11704].

A distinct continuous-time generalization appears in a scale-invariant future representation [1802.06426]. There, future outcomes are represented over a logarithmically compressed timeline, and multiplicative dilation in real time becomes an additive shift in log-time:
\[
\tau \mapsto a\tau \quad \Longleftrightarrow \quad s \mapsto s + \ln a.
\]
On a discrete logarithmic grid, moving to the next larger temporal scale is therefore an index shift. This is not a chunked or coarse-to-fine generative model, but it instantiates TENS in the precise sense that “next-scale” prediction corresponds to moving along a structured temporal axis rather than rolling out one step at a time.

These formulations differ in target semantics—discounted returns, future chunks, residual token maps, occupancy scales, or log-time future timelines—but they share the same organizing principle: prediction is distributed over a set of explicitly parameterized temporal scales, and adjacent scales are allowed to inform one another.

## 3. Learning mechanisms and architectural patterns

In the reinforcement-learning setting, TENS is implemented with standard TD\((\lambda)\) and linear function approximation [1112.1133]. With shared sparse tile-coded features \(\boldsymbol{\phi}(s_t) \in \{0,1\}^{6065}\), each predictor uses
\[
v^i_\gamma(s_t) \approx \mathbf{w}^i_\gamma{}^\top \boldsymbol{\phi}(s_t),
\]
\[
\delta_t^i = r_{t+1}^i + \gamma \,\mathbf{w}^i_\gamma{}^\top \boldsymbol{\phi}(s_{t+1}) - \mathbf{w}^i_\gamma{}^\top \boldsymbol{\phi}(s_t),
\]
\[
\mathbf{e}_t^i = \gamma \lambda \,\mathbf{e}_{t-1}^i + \boldsymbol{\phi}(s_t),
\]
and
\[
\mathbf{w}^i_\gamma \leftarrow \mathbf{w}^i_\gamma + \alpha\, \delta_t^i\, \mathbf{e}_t^i.
\]
The reported implementation used \(\lambda=0.9\), \(\alpha = 0.1/457\), and exactly \(457\) active binary features per time step. Practical efficiency came from sparse dot products, step-size scaling by the number of active features, multithreading, and per-\(\gamma\) trace sharing [1112.1133].

Later TENS systems replace value-function banks with hierarchical generative architectures. Next Forcing augments a flow-matching autoregressive video transformer with \(D=3\) lightweight MCP modules for depths \(d=1,2,3\), each receiving noisy target tokens for a future chunk and the previous depth’s representation [2606.11187]. Hidden states from layers \(\{4,12,20,30\}\) are fused by a two-layer MLP into \(h_{\mathrm{fuse}}\), and the per-depth fused input is
\[
z^{(d)} = W_d [h^{(d-1)}_{\mathrm{prev}}; Embed(x_{t_d}^{[d]})].
\]
This yields explicit causal chaining across temporal depths, with gradients from the multi-chunk losses flowing back into multiple backbone layers.

VideoAR uses a 3D multi-scale tokenizer and a causal Transformer backbone with block-wise causal masking to combine next-frame prediction and intra-frame next-scale prediction [2601.05966]. ScaleMoGen and MoScale transpose the same principle to motion: both generate coarse global structure first, then finer temporal detail, using hierarchical tokenization and shared or scale-conditioned autoregressive predictors [2605.11704; 2604.03799]. OccTENS applies the pattern to occupancy world modeling by splitting the problem into temporal scene-by-scene prediction and spatial scale-by-scale generation within each frame, implemented by a TensFormer with scale-wise temporal causal attention and frame-wise spatial attention [2509.03887].

A recurrent architectural pattern is therefore visible across domains. Coarser scales encode global structure, finer scales encode residual refinement, and cross-scale coupling is deliberately asymmetric: later predictions depend on earlier scales, but not conversely. This suggests that TENS is best understood as a causal ordering principle as much as a loss design.

## 4. Representative realizations across domains

The following examples illustrate how the same temporal next-scale idea has been instantiated in substantially different tasks.

| Paper | Prediction unit | Temporal next-scale structure |
|---|---|---|
| "Multi-timescale Nexting in a Reinforcement Learning Robot" [1112.1133] | Sensor and derived signals | Parallel value functions at \(\gamma \in \{0, 0.8, 0.95, 0.9875\}\) |
| "Next Forcing: Causal World Modeling with Multi-Chunk Prediction" [2606.11187] | Video chunks | next\(^1\), next\(^2\), next\(^3\) with causal chain |
| "VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction" [2601.05966] | Frame-scale residual maps | Frame-by-frame plus coarse-to-fine intra-frame generation |
| "ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation" [2605.11704] | Skeletal-temporal token maps | Coarse-to-fine scales preserving skeletal hierarchy |
| "Next-Scale Autoregressive Models for Text-to-Motion Generation" [2604.03799] | Motion token groups at temporal scales | Full-horizon coarse-to-fine temporal resolutions |
| "OccTENS: 3D Occupancy World Model via Temporal Next-Scale Prediction" [2509.03887] | Multi-scale occupancy tokens | Scene-by-scene temporal prediction and scale-by-scale spatial generation |

In robotics, the emphasis was on real-time anticipatory knowledge. The Critterbot system learned 2160 predictions over \(0.1\)–\(8\) s horizons, updated online at better than \(10\) Hz, with measured compute time of approximately \(55\) ms per \(100\) ms cycle and memory usage of approximately \(400\) MB [1112.1133]. In forecasting, the same broad multi-time-scale intuition appears in wind power prediction, where a TFT-based framework targets ultra-short-term, short-term, and medium-term horizons and reports nMAE reductions of \(31.75\%\) and \(20.79\%\) for 24-hour and 48-hour forecasting relative to the second-best model [2302.01222]. That paper does not introduce explicit hierarchical couplings across horizons; the data explicitly state that cross-scale information transfer occurs implicitly via decomposition and shared temporal processing. A plausible implication is that not all multi-horizon forecasting should be classified as TENS in the strong causal-coupling sense used by later generative papers.

In generative world models, the emphasis shifts to dense temporal supervision and efficient inference. Next Forcing reports a \(93.1\%\) relative improvement over LingBot-VA at \(5\)k training steps on RoboTwin at \(50\) fps, \(2.3\times\) faster convergence, and near-\(2\times\) inference throughput when the depth-1 head is retained at inference [2606.11187]. OccTENS reports mIoU \(22.06\) and IoU \(31.03\) with occupancy input on nuScenes val, and measured latency \(0.56\) s for the 6-scale model, compared with \(0.35\) s for OccWorld and approximately \(20\) s for OccSora [2509.03887]. In text-to-motion, ScaleMoGen reports FID \(0.030\) on HumanML3D and CLIP Score \(0.693\) on SnapMoGen, while MoScale reports strong alignment metrics on HumanML3D and KIT-ML with a coarse-to-fine temporal hierarchy [2605.11704; 2604.03799].

Taken together, these results show that TENS is not confined to one modality. It has been used for sensorimotor prediction, video generation, motion generation, occupancy world modeling, and multi-horizon forecasting, but the exact interpretation of “scale” varies from discount-controlled horizon to chunk depth, residual resolution, or hierarchical state granularity.

## 5. Computational properties, advantages, and trade-offs

One recurring motivation for TENS is efficiency relative to flat sequential prediction. In multi-timescale nexting, sparse shared features made thousands of online predictors practical: with \(A=457\) active features and per-\(\gamma\) eligibility sharing, complexity depended on \(\mathcal{O}(A\cdot P)\) for updates plus \(\mathcal{O}(F \cdot |\Gamma|)\) for trace maintenance, rather than dense \(\mathcal{O}(F\cdot P)\) processing [1112.1133]. In scale-invariant future estimation, the entire future timeline is constructed in one parallel operation over a logarithmic temporal grid, avoiding forward simulation whose cost grows linearly with horizon [1802.06426].

In chunked and scale-wise generative models, the claimed advantage is not merely lower wall-clock cost but improved supervision geometry. Next Forcing argues that next-chunk-only training is myopic; its MCP losses inject dense multi-scale temporal supervision into several backbone layers through fused intermediate features [2606.11187]. VideoAR argues that raster next-token autoregression over flattened tokens is poorly aligned with spatio-temporal structure, whereas next-scale generation reduces sequence length and better captures spatial correlations [2601.05966]. MoScale makes the same argument for motion: local continuity makes next-token prediction easy, but it does not force commitment to global semantics such as counts, ordering of actions, and global trajectories [2604.03799].

These advantages come with clear trade-offs. In discounted-return TENS, higher \(\gamma\) extends horizon but increases variance and can slow learning; lower \(\gamma\) learns faster but misses longer dependencies [1112.1133]. In coarse-to-fine generators, mistakes at coarse scales can propagate to fine scales. Several papers therefore introduce explicit correction mechanisms. MoScale perturbs coarser tokens during training through cross-scale hierarchical refinement and uses in-scale temporal refinement for selective bidirectional re-prediction [2604.03799]. VideoAR uses Cross-Frame Error Correction and Random Frame Mask to mitigate exposure bias and over-reliance on distant frames [2601.05966]. Next Forcing keeps depth-2 and depth-3 predictions for training but uses only depth-1 at inference to avoid drift [2606.11187].

A common misconception is that TENS is simply multi-horizon prediction. The cited works suggest a narrower technical reading. TENS generally involves explicit organization of prediction across adjacent temporal scales, together with some mechanism by which one scale conditions, constrains, or refines another. Mere production of several horizons from a shared backbone, without such cross-scale structure, is closer to generic multi-horizon forecasting than to the stronger TENS formulations.

## 6. Generalizations, limitations, and research directions

The literature identifies several extensions that broaden TENS beyond its initial forms. In reinforcement learning, state-dependent discounting replaces constant \(\gamma\) with variable \(\gamma_t\), allowing event-conditioned horizons such as predicting power expenditure until a light sensor saturates or until a stochastic termination [1112.1133]. Off-policy prediction can be handled with gradient TD methods such as GTD2 and TDC, while non-linear representations such as kernels, random Fourier features, and neural networks may reduce residual error when linear tile coding underfits [1112.1133]. Continuous-time formulations based on Laplace transforms and the Post approximation offer a scale-invariant future timeline retaining “what will happen when,” rather than only a scalar discounted value [1802.06426].

In generative modeling, later work pushes TENS toward stronger causal world modeling and controllability. OccTENS integrates ego-motion as a \(0\)-th scale token and treats occupancy and motion in one autoregressive sequence, enabling controllable future occupancy generation aligned with planned pose tokens [2509.03887]. MSTP with IG-MC extends the next-scale concept along both temporal and state axes, using synchronized multi-scale visual previews and a multi-agent hierarchy for general and surgical scenes [2509.17429]. This suggests that the “scale” in TENS need not be purely temporal resolution; it can also denote hierarchical state abstraction, provided that cross-scale consistency is maintained.

The limitations are similarly recurrent. Performance is bounded by representational capacity in linear nexting, by hierarchy design in motion tokenization, and by memory or compute overhead in multi-head or multi-agent systems [1112.1133; 2605.11704; 2509.17429]. Very long horizons remain difficult. Next Forcing notes that depth-2 and depth-3 are excellent for training but risky for inference because far-future chunk prediction can accumulate drift [2606.11187]. OccTENS notes that increasing the number of scales improves fidelity but raises latency, and its motion tokenizer omits the \(z\)-axis, which is reasonable for on-road driving but limits generalization [2509.03887]. MoScale reports that excessive corruption or too many refinement iterations can degrade stability [2604.03799].

The available evidence therefore supports a precise characterization of TENS. It is neither a single algorithm nor a synonym for forecasting. It is a design paradigm in which temporal prediction is decomposed into an ordered family of next-scales—discount horizons, chunk depths, coarse-to-fine resolutions, or logarithmic future lags—and these scales are tied together by shared representations, causal coupling, or refinement operators. The historical trajectory from multi-timescale nexting to modern generative world models suggests that TENS has evolved from a bank of short-horizon predictors into a broader framework for organizing temporal structure in learning systems [1112.1133; 2606.11187].

Source: https://www.emergentmind.com/topics/temporal-next-scale-prediction-tens