---
title: 'TConstFormer: Constant-Time Transformer'
url: https://www.emergentmind.com/topics/tconstformer
type: topic
---

# TConstFormer: Constant-Time Transformer

Searching arXiv for the specified paper and closely related work to ground the article.
arxiv.search code:
{"query":"id:2509.00202 OR title:\"From TLinFormer to TConstFormer\"","max_results":5,"sort_by":"submittedDate","sort_order":"descending"}
arxiv.search code:
{"query":"TConstFormer TLinFormer","max_results":10,"sort_by":"relevance","sort_order":"descending"}
TConstFormer is a Transformer architecture introduced in "From TLinFormer to TConstFormer: The Leap to Constant-Time Transformer Attention: Achieving O(1) Computation and O(1) KV Cache during Autoregressive Inference" [2509.00202]. It is designed for autoregressive inference on ultra-long or effectively unbounded sequences, where standard decoder-only Transformers suffer from a linearly growing KV cache and sequence-length-dependent latency. The architecture builds on TLinFormer, but replaces direct access from current generation units to the full raw history with a fixed-size encoded context state. Its defining claims are a KV cache whose footprint is $\mathcal{O}(1)$ with respect to total processed sequence length and a periodic state update mechanism that yields constant-time computation on cache hits and amortized $\mathcal{O}(1)$ behavior across long streams.

## 1. Motivation and predecessor architecture

The central scaling problem addressed by TConstFormer is standard autoregressive attention. In a decoder-only Transformer, each new token attends to all previous tokens, and the paper writes the attention operator as
\[
\text{Attention}(\mathbf{Q}, \mathbf{K}, \mathbf{V})
=
\text{softmax}\left(\frac{\mathbf{Q}\mathbf{K}^\top}{\sqrt{d_k}}\right)\mathbf{V}.
\]
If $\mathbf{Q}, \mathbf{K}, \mathbf{V} \in \mathbb{R}^{L \times d}$, computing $\mathbf{Q}\mathbf{K}^\top$ costs $\mathcal{O}(L^2 d)$, and the full-sequence complexity per layer is described as $\mathcal{O}(N^2 d)$. During autoregressive inference, the model avoids recomputing all past attention by storing Keys and Values for all previous tokens, so the KV cache grows linearly with sequence length. The paper gives the per-session cache formula
\[
M_{\text{transformer}} =
2 \cdot B \cdot L \cdot d_{\text{model}} \cdot P_{\text{bytes}} \cdot N_{\text{layers}},
\]
which makes the memory cost $\mathcal{O}(N d)$ [2509.00202].

This scaling is problematic for ultra-long sequences and streaming settings. The paper identifies two practical failure modes: memory exhaustion from linearly increasing cache size, and rising latency because per-token compute and memory bandwidth costs still grow with $N$. It also positions windowed or sliding attention as an incomplete remedy: such methods bound compute and memory by truncating history, but at the cost of losing distant context and inducing what the paper calls “catastrophic forgetting” of long-range dependencies.

TConstFormer is presented as the successor to TLinFormer. TLinFormer reinterprets attention as a fully connected layer over the sequence-length dimension, or “L-dimension MLP,” and rewires connectivity so that attention can be evaluated exactly, rather than via kernel approximation, with linear complexity $\mathcal{O}(N d)$. However, TLinFormer still depends on explicit access to the full historical sequence and therefore retains a sequence-length-dependent $\mathcal{O}(N d)$ KV cache. TConstFormer is introduced specifically to remove that remaining memory dependence.

## 2. Architectural organization

TConstFormer preserves TLinFormer’s attention-building viewpoint but changes the topology of information flow. The input is partitioned into a historical context window $X_{\text{hist}}$ and a generation window $X_{\text{gen}}$. Each TConstFormer block then has two coupled paths: a context path for the historical segment and a generation path for the currently generated segment [2509.00202].

In the context path, the first layer compresses the historical context along the sequence-length dimension using Focused Attention. Middle layers apply self-attention within the compressed context representation. A final layer may expand along the sequence dimension using Cross Attention if the block is stacked. In the generation path, each internal layer combines two operations: causal self-attention within $X_{\text{gen}}$, which the paper terms “Internal Cohesion,” and cross-attention from the generation window to the context representation, which it terms “Context Integration.” Their outputs are combined and passed through an FFN to produce the next hidden state.

The decisive architectural change is that direct connections from raw historical inputs to the current generation units are removed. Historical information must first pass through the compressed context representation, and only that representation is available to the generation path. The paper states that this interruption of direct information pathways makes the inference state depend on a fixed-size hidden state rather than on the entire growing historical sequence. In effect, TConstFormer converts the historical side of the model into a bounded-state encoder while preserving Transformer-style modular attention blocks on both the context and generation sides.

The paper frames this as a topological rather than merely algorithmic modification. TLinFormer keeps full-history connectivity but reduces computational complexity; TConstFormer changes the graph of accessible dependencies so that the model no longer needs to store the raw token-by-token past in its inference state.

## 3. Constant-state cache and periodic synchronization

The operational mechanism that yields constant-state inference is windowed processing with periodic context updates. The model uses a historical context window length $W_{oh}$ and a generation window length $W_{og}$, with total observation window
\[
W_{\text{total}} = W_{oh} + W_{og}.
\]
During inference, the system stores only the compressed representation of the historical context window and the representation of the current generation window. New tokens are generated one by one within the generation window; once $W_{og}$ new tokens have been produced, the window slides and a global synchronization step recomputes or updates the context path from the enlarged history [2509.00202].

The cache formula given in the paper is
\[
M_{\text{TConstFormer}} =
2 \cdot B \cdot (H + 1) \cdot W_{oh} \cdot d_{\text{model}}
+
2 \cdot B \cdot (H + 2) \cdot W_{og} \cdot d_{\text{model}}.
\]
Because $W_{oh}$, $W_{og}$, $H$, $d_{\text{model}}$, and $B$ are fixed or bounded after training, the cache footprint has no dependence on total processed sequence length $N$. This is the basis for the paper’s claim that the KV cache is $\mathcal{O}(1)$ with respect to $N$.

The same section distinguishes cache misses from cache hits. On a cache miss, corresponding to training or to a global synchronization step, the context path must be recomputed. The paper gives the total cost as
\[
\text{Total Cost}
=
D \left[
N(2W_{oh})
+
H(W_{oh}^2 + W_{og}^2 + W_{oh}W_{og})
+
2W_{og}^2 - W_{og}W_{oh}
\right],
\]
and then simplifies it to
\[
\text{Total Cost} = C_1 \cdot N + C_0,
\]
so the miss complexity is $\mathcal{O}(N)$.

On a cache hit, the context path is entirely reused and only the generation window is active. The paper gives the constant-cost expression
\[
\text{Total Cost}
=
(H + 1) D W_{oh} + (H + 2) D W_{og}^2,
\]
which contains no $N$. This is the formal basis for constant-time cache-hit inference.

| Regime | Compute characterization | KV cache characterization |
|---|---|---|
| Standard Transformer | $\mathcal{O}(N^2 d)$ per layer for full sequence | $\mathcal{O}(N d)$ |
| TLinFormer | $\mathcal{O}(N d)$ | $\mathcal{O}(N d)$ |
| TConstFormer | Cache hit: constant in $N$; cache miss: $\mathcal{O}(N)$ | $\mathcal{O}(1)$ in $N$ |

A common misunderstanding is to collapse these two regimes into a single unconditional constant-time statement. The paper’s formulation is more specific: constant-time behavior applies to cache hits within a fixed generation window, whereas periodic synchronization remains linear in the observed history. The broader claim is amortized $\mathcal{O}(1)$ behavior over long streams because the expensive update occurs only every $k = W_{og}$ steps.

## 4. Mathematical formulation, training procedure, and inference workflow

The paper uses a generic tensorized attention definition with
\[
\mathbf{Q} \in \mathbb{R}^{B \times L_Q \times D}, \qquad
\mathbf{K}, \mathbf{V} \in \mathbb{R}^{B \times L_K \times D},
\]
and interprets attention as a fully connected transformation along the length dimension. This “dimension transformation” view motivates its treatment of self-attention, causal self-attention, focused attention, and cross-attention as different connectivity patterns over the sequence axis [2509.00202].

For cache misses, the paper decomposes cost into a left window and right window. On the historical side, the first cross-attention from context window to full history costs $D (N - W_{og}) W_{oh}$, the $H$ self-attention layers within the context window cost $H D W_{oh}^2$, and the final cross-attention used to restore length costs another $D (N - W_{og}) W_{oh}$. This gives
\[
C_{\text{left}} = 2D(N - W_{og})W_{oh} + HDW_{oh}^2.
\]
On the generation side, the cross-attention from generation window to context window across $H+1$ layers costs $(H+1)DW_{oh}W_{og}$, and causal self-attention within the generation window across $H+2$ layers costs $(H+2)DW_{og}^2$, producing
\[
C_{\text{right}} = (H + 1)DW_{oh}W_{og} + (H+2)DW_{og}^2.
\]
The total miss cost is $T = C_{\text{left}} + C_{\text{right}}$.

Training is performed with a sliding-window chunking procedure. A long input sequence $X$ is split into chunks $X_0, X_1, X_2, \dots$, each of length $W_{og}$. Chunk $0$ is processed over $[0, W_{og}]$ with empty historical context. The window then slides with stride $S = W_{og}$: chunk $1$ processes $[W_{og}, 2W_{og}]$ using $[0, W_{og}]$ as historical context, chunk $2$ processes $[2W_{og}, 3W_{og}]$ using $[0, 2W_{og}]$ as historical context, and so on. Outputs from generation windows are concatenated for the final loss calculation.

This training protocol has two stated effects. First, it teaches the model to use a bounded observation window $W_{oh} + W_{og}$ to represent arbitrarily long history. Second, because the generation path is not allowed to access all raw historical tokens directly, it forces the context path to act as a learned encoding module for long history and the generation path to act as a decoder or generator conditioned on that compressed state. The loss remains standard language-model cross-entropy; the paper explicitly notes that there are no exotic auxiliary losses.

At inference time, the workflow mirrors the training design. An initial prompt is converted into context state through one cache-miss pass. Token generation inside the current window then proceeds with constant cache size and constant cache-hit cost. When the generation window fills, the model slides the window, adds those tokens to history, recomputes the context path, and resumes token-by-token generation.

## 5. Experimental configuration and empirical behavior

All experiments reported in the paper are conducted on wikitext-103-v1, described as approximately 120M tokens, using a single RTX 4090 GPU [2509.00202]. The model configuration is fixed across the comparison: vocabulary size 50,257, embedding dimension $n_{\text{embd}} = 432$, 12 attention heads, and total Transformer depth 8. The baseline is a standard decoder-only Transformer with 8 layers. TConstFormer uses 2 TConstFormer blocks, each with internal depth $H = 2$, yielding equivalent total depth of 8 and approximately 41M parameters. Training hyperparameters such as learning rate and batch size are kept identical to isolate architectural effects.

The reported training overhead is higher for TConstFormer because chunked sliding-window processing introduces extra scheduling and recomputation. At sequence length 1K, the baseline Base 1K requires roughly 620 seconds per epoch, whereas TConstFormer 1K-1K-0.5 requires roughly 890 seconds, corresponding to about 42% overhead.

Validation perplexity is reported for 512, 1K, and 2K settings. At 512 sequence length, Base 512 obtains 21.6, TLinFormer 512-512-* variants obtain 21.9, and TConstFormer 512-512-0.5 obtains 21.6. At 1K sequence length, Base 1K obtains 22.5, TLinFormer 1K-1K-0.5 obtains 22.7, and TConstFormer 1K-1K-0.5 obtains 22.7. At 2K sequence length, Base 2K obtains 29.5, TLinFormer 2K-2K-0.5 obtains 29.8, and TConstFormer 2K-2K-0.5 obtains 29.6. The paper therefore characterizes TConstFormer as matching baseline performance when its observation window length equals the baseline context length and as slightly outperforming TLinFormer in some settings.

The latency and memory experiments are more distinctive. The baseline Transformer shows super-linear, near-$\mathcal{O}(N^2)$ latency growth with sequence length, and its cache speedup factor peaks at approximately $1.26\times$ before decaying toward $1\times$ at large $N$, which the paper attributes to memory-bandwidth bottlenecks. TLinFormer exhibits a dual-mode latency profile: cache misses grow linearly with $N$, and cache hits remain on a lower-slope linear curve, with cache speedup greater than $10\times$ for million-token sequences. TConstFormer’s cache-hit latency curve is described as essentially flat, its cache speedup exceeds $40\times$ on million-token sequences, and its cache memory usage remains constant as $N$ increases. These empirical observations are presented as corroboration of the architecture’s $\mathcal{O}(1)$-style inference state and constant-time cache-hit regime.

## 6. Relation to efficient attention methods, recurrent-state models, and design trade-offs

The paper places TConstFormer within a broader family of efficient sequence models [2509.00202]. Relative to Longformer, the distinction is that Longformer uses sliding-window plus global-token attention and therefore truncates history beyond a fixed window, whereas TConstFormer keeps a compressed representation of the entire history visible through the context path. Relative to Linformer, the distinction is that Linformer reduces memory by projecting Keys and Values to a lower-rank space, which still scales with sequence length and uses approximation, whereas TConstFormer maintains a constant cache and uses exact attention over its own windows. Relative to Performer and other kernel-based linear attention methods, the difference is the absence of random-feature approximation. Relative to FlashAttention and related IO-aware implementations, the difference is architectural: FlashAttention changes memory access patterns, but not the fact that the inference state grows with sequence length. The paper also notes a conceptual resemblance to recurrent and state-space models such as RWKV and SSMs, since these maintain fixed-size hidden states, but argues that TConstFormer retains attention-like mechanisms and Transformer-like modularity.

These comparisons clarify an important point about the word “exact.” In the paper, exactness refers to the attention computation within the TConstFormer connectivity topology; it does not mean the architecture preserves the raw full-history token-to-token connectivity of a standard Transformer. The architectural intervention is precisely the removal of those direct links.

The design introduces explicit trade-offs. The hyperparameters $W_{oh}$ and $W_{og}$ determine latency, throughput, and compression pressure. Larger $W_{og}$ reduces the frequency of $\mathcal{O}(N)$ synchronization events and improves the amortized constant, but increases the constant per-token cache-hit cost because of the $W_{og}^2$ term in the hit complexity. Larger $W_{oh}$ increases the capacity of the compressed context state, whereas smaller $W_{oh}$ increases compression pressure and may slightly hurt quality on extremely long contexts. The paper reports that perplexity is relatively robust to moderate variations in the ratio $W_{oh}/W_{\text{total}}$, but that extreme compression can degrade performance.

A second trade-off concerns task type. The reported evaluation uses a model with approximately 41M parameters and Wikipedia-scale training data, not billion-parameter LLMs. The paper explicitly notes that TConstFormer’s constant-state representation appears well suited to macroscopic semantic patterns and generalization, but that it may be less suited to tasks requiring exact recall of very small details buried deep in the history. It also notes the absence of direct tests on long-context retrieval benchmarks such as Needle-in-a-Haystack.

## 7. Broader implications and prospective extensions

The paper presents TConstFormer as evidence that a Transformer-like architecture can be reconfigured so that its inference state remains finite and independent of total sequence length [2509.00202]. In the authors’ framing, this opens a route to streaming language models that can process unbounded token streams without ever accumulating a linearly growing KV cache. A plausible implication is that the architecture is particularly aligned with deployment regimes in which inference dominates training cost, since the reported training overhead is one-time while the serving benefits recur across long-running sessions.

Several future directions are named explicitly. One is integration with Mixture-of-Experts, where TConstFormer’s constant-time and constant-memory per-token behavior is presented as complementary to MoE parameter efficiency. Another is application to streaming modalities such as real-time video, continuous sensor data, and unending conversations, all of which fit the paper’s emphasis on online processing. A third is extension from one-dimensional sequences to higher-dimensional tensorial attention, for example video tensors of the form $[T, H, W, C]$, with the aim of learning constant-state representations across space and time.

The paper also advances a more general interpretation of “constant-state representation.” It suggests that an intelligent agent operating in an effectively infinite world must nonetheless maintain a finite internal state, and treats TConstFormer as a step toward architectures that are forced to learn scale-invariant abstraction and compression rather than relying on raw sequence length as an auxiliary resource. This is a conceptual extrapolation rather than an experimentally established property, but it situates the model within a broader line of inquiry concerning bounded-state reasoning in continuous environments.

Source: https://www.emergentmind.com/topics/tconstformer