---
title: KV-Invariant Transformer Expansion
url: https://www.emergentmind.com/topics/kv-invariant-transformer-expansion-kite
type: topic
---

# KV-Invariant Transformer Expansion

KV-Invariant Transformer Expansion (KITE) is a model-scaling paradigm introduced for reducing the computational cost of training, prompt prefilling, and autoregressive inference while increasing model capacity. KITE expands a smaller Transformer by placing newly added parameters in a region that reads, but does not produce, the reusable attention key–value (KV) cache. Its concrete instantiation, the Step Scale Transformer (SST), is a two-tower decoder comprising a KV-producing **Prefiller** and a KV-reading **Decoder**. In bulk prompt processing, only the Prefiller processes all prompt tokens and constructs the KV cache; during generation, both towers process each newly generated token. A 66.959B-parameter SST model with 2.155B active body parameters per decode token achieved lower training loss than 47B and 63B MoE Transformer baselines at comparable cumulative theoretical training compute, while its analytical inference-cost proxy was 6.7% lower than the 47B model and 31.6% lower than the 63B model [2609.27294].

## 1. Conceptual basis and computational objective

KITE separates Transformer capacity into two computationally distinct regions: a **KV-producing region**, whose operations determine the reusable attention memory, and a **KV-reading region**, whose operations consume that memory without generating KV representations required by future positions. The paradigm is motivated by the observation that training, prompt prefilling, and autoregressive decoding impose different computational requirements.

During training, both the KV-producing and KV-reading regions are evaluated at every supervised position so that the newly added parameters can be optimized. During bulk prefilling, however, the KV-reading region need not process every prompt position if its computation is token-local once the corresponding Prefiller outputs and KV cache are available. During autoregressive decoding, both regions remain active because each generated token must extend the Prefiller cache and be processed by the Decoder.

The central invariant is computational and causal rather than parameter freezing or exact functional equivalence. Let $P$ denote the Prefiller and $D$ the Decoder:

$$
(M_t,z_t)=P(x_{1:t};\theta_P),
$$

$$
\ell_t=D(x_t,z_t,M_t;\theta_D),
$$

where $M_t$ is reusable attention memory, principally the KV cache, $z_t$ denotes optional Prefiller features, and $\ell_t$ denotes the next-token logits. Future positions’ reusable memory is produced only by the Prefiller, while the Decoder reads the Prefiller-produced memory and does not generate KV required by later positions.

This construction differs from simply enlarging a conventional Transformer. In an ordinary causal Transformer, all layers process all prompt tokens during prefilling, and increasing model depth, width, or active capacity generally increases prompt-wide computation. KITE instead expands the KV-reading path while retaining a smaller KV-producing path.

## 2. Step Scale Transformer architecture

The Step Scale Transformer (SST) implements KITE with two same-depth Transformer towers:

- **Prefiller**: produces layer-wise keys and values.
- **Decoder**: produces queries, reads the Prefiller’s layer-wise keys and values, and generates predictions.

In the reported configuration, each tower contains 18 layers, giving an effective 18-layer Prefiller and 18-layer Decoder. The towers use hidden width 2304, 512 routed experts, Top-8 routing, and 16 MoE layers; the detailed architecture begins with two dense feed-forward layers. The attention pattern is denoted SSSF, consisting of three sliding-window layers followed by one full-attention layer, with the stated remainder pattern for the 18-layer configuration.

For Decoder layer $l$ and token position $t$, the query is generated from the Decoder hidden state:

$$
q^D_{l,t}=Q^D_l(h^D_{l-1,t}).
$$

The Decoder then attends to the corresponding Prefiller representations:

$$
a^D_{l,t}
=
\operatorname{Attn}
\left(
q^D_{l,t},
K^P_l[\mathcal I_l(t)],
V^P_l[\mathcal I_l(t)]
\right),
$$

where $\mathcal I_l(t)$ is the set of positions permitted by the causal or sliding-window mask. The Prefiller produces $K^P_l$ and $V^P_l$, whereas the Decoder produces queries and consumes those KV tensors. The effective architecture contains K/V projections and K normalization in the Prefiller; the Decoder has independent residual, FFN/MoE, and query-side parameters but does not produce an additional cache required by later positions.

The Decoder is initialized through a parameter-free RMS-normalized bridge:

$$
h^D_{0,t}
=
\operatorname{RMS}_{\epsilon}(e_t)
+
\operatorname{RMS}_{\epsilon}(h^P_{L,t}),
$$

with

$$
\operatorname{RMS}_{\epsilon}(v)
=
\frac{v}{\sqrt{\operatorname{mean}(v^2)+\epsilon}},
\qquad
\epsilon=10^{-5}.
$$

The final Decoder state is passed through the shared output normalization and untied output head. The bridge and final-state RMS operations have no trainable parameters, accumulate in FP32, and return the input dtype in the reported implementation.

## 3. Execution across training, prefilling, and decoding

### Training

After expansion, both towers process every supervised position. The Prefiller processes the sequence and generates layer-wise $K^P,V^P$ tensors. The Decoder computes predictions at all training positions, and gradients flow through its attention reads back into the Prefiller through the reused KV tensors. The Prefiller is therefore jointly trained after expansion rather than frozen.

For source and continuation token counts $D_s,D_c$, and forward FLOPs per token $f_s,f_c$, the paper estimates training cost as

$$
C_{\mathrm{train}}
=
3(D_s f_s+D_c f_c),
$$

where the factor 3 approximates forward plus backward computation. The theoretical accounting excludes optimizer cost, communication, checkpoint I/O, rematerialization, conversion overhead, normalization and residual micro-operations, and auxiliary loss terms.

### Bulk prompt prefilling

For a prompt $x_{1:n}$, the Prefiller processes all $n$ tokens and constructs the layer-wise cache

$$
\mathcal M_n
=
\{(K^P_l[1:n],V^P_l[1:n])\}_l.
$$

The Decoder is not evaluated at every prompt position. It is evaluated only at the final prompt position $n$, using the Prefiller KV cache to produce the first output-token distribution. Earlier Decoder computations can be omitted because the Decoder does not generate memory needed by later prompt positions. The bulk-prefill path is therefore source-sized apart from the final-position Decoder computation and shared output operations.

### Autoregressive decoding

For each generated token, the Prefiller processes the token and appends its K/V representations to the cache. The Decoder then processes that token, reads the expanded KV cache, and produces the next-token logits. KITE therefore does not eliminate the additional Decoder computation during generation. Its advantage depends on amortizing this decode overhead over sufficiently substantial prompt-prefill work.

## 4. Upcycling and parameter construction

The evaluated SST model was expanded from a 33.819B-parameter source model. The construction procedure retained the source Transformer as the SST Prefiller, copied each source layer into the corresponding Decoder layer, shared the embedding, output normalization, and output head, and then continued training both towers jointly.

This is aligned source-to-target expansion rather than strict function-preserving Net2Net expansion. Adding the Decoder changes the prediction graph, so the expanded model experiences a loss increase at conversion before adapting during continuation training. The Prefiller and shared-parameter optimizer states are retained, while fresh optimizer states are initialized for the Decoder. Training continues on the inherited cosine learning-rate schedule without a separate Decoder warmup.

The source stage used 167.98B tokens and the continuation stage used 222.55B tokens, for 390.54B total tokens at the evaluated SST endpoint. Training used BF16, Muon and Adam parameter groups, 128 NVIDIA H800 GPUs, expert parallelism of 8, sequence length 4096, and a global batch size of 3072 sequences. The continuation stage comprised 17,687 updates after 13,350 source updates.

| Model | Total parameters | Layers | Decode-active body parameters per token |
|---|---:|---:|---:|
| Source | 33.819B | 18 | 1.120B |
| 67B SST | 66.959B | $18+18$ | 2.155B |
| 47B classic | 46.727B | 20 | 1.477B |
| 63B classic | 62.691B | 22 | 2.016B |

The SST model contains 66.365B body parameters and 0.594B embedding-plus-output-head parameters. Its bulk-prefill active body count is 1.120B, corresponding to the Prefiller-only path. The active counts include selected Top-8 experts, shared experts, routers, and other body weights required for one token, while excluding the embedding and output head.

## 5. Training quality and scaling results

The 47B classic model defines the normalized theoretical training budget:

$$
C_{\mathrm{47B}}
\approx
5.16836\times 10^{21}\ \text{FLOPs}.
$$

At their reported final endpoints, the 47B classic, 63B classic, and 67B SST models each use approximately the same relative theoretical training FLOPs. Their training-token counts are 443.16B, 334.62B, and 390.54B, respectively.

The final EMA-200 training losses were:

| Model | Training tokens | Final training loss |
|---|---:|---:|
| 47B classic | 443.16B | 1.6006 |
| 63B classic | 334.62B | 1.5921 |
| 67B SST | 390.54B | 1.5900 |

SST’s final loss was 0.0106 lower than the 47B baseline and 0.0021 lower than the 63B baseline at comparable cumulative theoretical training compute. Its loss increased at the expansion point and subsequently decreased during continuation training.

The reported downstream results were:

| Task | 47B classic | 63B classic | 67B SST |
|---|---:|---:|---:|
| OpenBookQA | 68.50 | 71.00 | 75.00 |
| MMLU | 58.44 | 59.39 | 60.54 |
| GSM8K | 56.56 | 55.04 | 58.38 |
| MATH | 33.06 | 33.14 | 34.36 |
| HumanEval, few-shot | 35.98 | 32.93 | 37.80 |
| MBPP, 3-shot | 50.40 | 50.80 | 51.40 |
| BBH | 50.36 | 52.56 | 52.74 |
| ARXIV NLL | 1.4521 | 1.4424 | 1.4421 |

The comparisons used the same tokenizer, data recipe, training protocol, and evaluation checkpoints. The reported results support the claim that the expanded two-tower model can improve quality at comparable cumulative theoretical training compute, although the experiments do not establish a general scaling law or isolate the causal contribution of each architectural component.

## 6. Inference-cost accounting, workload dependence, and limitations

KITE’s inference-cost advantage is workload-dependent. The paper uses an analytical active-parameter proxy rather than measured latency, throughput, energy, or serving cost. Let $p$ denote bulk-prefill active-body parameters normalized to the 47B classic model, $d$ denote decode-active parameters under the same normalization, and $w$ denote the prefill cost weight:

$$
C_{\mathrm{proxy}}(w)
=
wp+(1-w)d.
$$

For the illustrative 75:25 prefill-to-decode mixture,

$$
C_{\mathrm{proxy}}
=
0.75p+0.25d.
$$

The normalized proxy is 1 for the 47B classic model, 0.933 for SST, and 1.365 for the 63B classic model. SST is consequently estimated to be 6.7% cheaper than the 47B model and 31.6% cheaper than the 63B model under this proxy.

The reported break-even points are approximately:

- **Against 63B classic**: 13.4% prefill and 86.6% decode.
- **Against 47B classic**: 65.5% prefill and 34.5% decode.

Thus SST is estimated to be cheaper than the 47B model when prefill constitutes more than roughly 65.5% of weighted cost, and cheaper than the 63B model for much less input-heavy mixtures. Applying selected OpenRouter traffic charge shares of 67.4%–76.7% as proxy prefill weights gives estimated SST savings of 1.3%–7.9% versus the 47B model and 27.7%–32.5% versus the 63B model. These traffic statistics are price-based workload indicators rather than direct measurements of GPU prefill work, and the selected traffic is not labeled specifically as agentic usage.

KITE is particularly suited to workloads involving repeated long-context processing, including repeated agent turns, tool-use loops, accumulated conversation state, large documents, and large code contexts. In short-prompt or decode-dominated workloads, the additional Decoder computation may eliminate the advantage.

The architecture does not eliminate the Prefiller’s KV cache, reduce its number of cached tokens, or directly replace MQA, GQA, cross-layer attention, or KV-compression methods. The paper does not provide a complete byte-level KV-memory comparison, measured serving speedups, or measured maximum-context improvements. Running two towers may also introduce synchronization, memory-bandwidth, expert-routing, and parallelization overheads not captured by the active-parameter proxy.

A further limitation is component attribution. SST combines upcycling, two-tower execution, MoE routing, layer-wise KV reuse, and the specified attention pattern. The experiments do not determine how much of the observed improvement arises specifically from KV invariance, initialization, the source model, or other architectural choices. The results therefore establish an empirical instance of KITE rather than a complete theory of Transformer expansion.

KITE’s principal distinction from KV-only self-attention is architectural. Key-Value Transformer variants remove the independent query projection from self-attention and replace $QK^{\mathsf T}$ with $KK^{\mathsf T}$, yielding a lower-cost attention mechanism with symmetric pre-softmax affinities [2305.19129]. KITE does not remove queries from the Decoder or alter the self-attention similarity rule in that manner. Instead, it preserves a source-sized KV-producing path and adds a separate KV-reading capacity-expansion path. Its central claim is consequently about where additional Transformer computation is placed, not about replacing QKV attention within an individual attention block.

Source: https://www.emergentmind.com/topics/kv-invariant-transformer-expansion-kite