Papers
Topics
Authors
Recent
Search
2000 character limit reached

KV-Invariant Transformer Expansion

Updated 30 September 2026
  • KV-Invariant Transformer Expansion (KITE) separes Transformer computations into two regions: a KV-producing region for long-term memory storage and a KV-reading region which reduces computational cost for tasks like prompt prefilling and autoregressive decoding, making for optimal training and inference performance while reducing the computational requirement. .
  • The Step Scale Transformer (SST), implementing the KITE paradigm, consists of a Prefiller that generates KV representations and a Decoder that reads these representations, allowing efficient model scaling and parallelism.
  • An experiment showed that a 67B SST model lowered training loss and inferred cost compared to 47B and 63B comparator models, and improved downstream task performance.].
  • follow_up_questions
  • How effective is the Step Scale Transformer (SST) in handling variable prompt lengths and dynamic contexts?
  • In what hardware configurations and runtime environments is KITE most beneficial, and how does its efficiency vary? F
  • What architectural modifications improve the efficiency of KV-cache reuse in the Step Scale Transformer?
  • is the performance gain of KITE consistent across different datasets and tasks?--, Find recent papers about Model Scaling Improvements.

KV-Invariant Transformer Expansion (KITE) is a model-scaling paradigm introduced for reducing the computational cost of training, prompt prefilling, and autoregressive inference while increasing model capacity. KITE expands a smaller Transformer by placing newly added parameters in a region that reads, but does not produce, the reusable attention key–value (KV) cache. Its concrete instantiation, the Step Scale Transformer (SST), is a two-tower decoder comprising a KV-producing Prefiller and a KV-reading Decoder. In bulk prompt processing, only the Prefiller processes all prompt tokens and constructs the KV cache; during generation, both towers process each newly generated token. A 66.959B-parameter SST model with 2.155B active body parameters per decode token achieved lower training loss than 47B and 63B MoE Transformer baselines at comparable cumulative theoretical training compute, while its analytical inference-cost proxy was 6.7% lower than the 47B model and 31.6% lower than the 63B model (Hu et al., 23 Sep 2026).

1. Conceptual basis and computational objective

KITE separates Transformer capacity into two computationally distinct regions: a KV-producing region, whose operations determine the reusable attention memory, and a KV-reading region, whose operations consume that memory without generating KV representations required by future positions. The paradigm is motivated by the observation that training, prompt prefilling, and autoregressive decoding impose different computational requirements.

During training, both the KV-producing and KV-reading regions are evaluated at every supervised position so that the newly added parameters can be optimized. During bulk prefilling, however, the KV-reading region need not process every prompt position if its computation is token-local once the corresponding Prefiller outputs and KV cache are available. During autoregressive decoding, both regions remain active because each generated token must extend the Prefiller cache and be processed by the Decoder.

The central invariant is computational and causal rather than parameter freezing or exact functional equivalence. Let PP denote the Prefiller and DD the Decoder:

(Mt,zt)=P(x1:t;θP),(M_t,z_t)=P(x_{1:t};\theta_P),

ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),

where MtM_t is reusable attention memory, principally the KV cache, ztz_t denotes optional Prefiller features, and ℓt\ell_t denotes the next-token logits. Future positions’ reusable memory is produced only by the Prefiller, while the Decoder reads the Prefiller-produced memory and does not generate KV required by later positions.

This construction differs from simply enlarging a conventional Transformer. In an ordinary causal Transformer, all layers process all prompt tokens during prefilling, and increasing model depth, width, or active capacity generally increases prompt-wide computation. KITE instead expands the KV-reading path while retaining a smaller KV-producing path.

2. Step Scale Transformer architecture

The Step Scale Transformer (SST) implements KITE with two same-depth Transformer towers:

  • Prefiller: produces layer-wise keys and values.
  • Decoder: produces queries, reads the Prefiller’s layer-wise keys and values, and generates predictions.

In the reported configuration, each tower contains 18 layers, giving an effective 18-layer Prefiller and 18-layer Decoder. The towers use hidden width 2304, 512 routed experts, Top-8 routing, and 16 MoE layers; the detailed architecture begins with two dense feed-forward layers. The attention pattern is denoted SSSF, consisting of three sliding-window layers followed by one full-attention layer, with the stated remainder pattern for the 18-layer configuration.

For Decoder layer ll and token position tt, the query is generated from the Decoder hidden state:

ql,tD=QlD(hl−1,tD).q^D_{l,t}=Q^D_l(h^D_{l-1,t}).

The Decoder then attends to the corresponding Prefiller representations:

DD0

where DD1 is the set of positions permitted by the causal or sliding-window mask. The Prefiller produces DD2 and DD3, whereas the Decoder produces queries and consumes those KV tensors. The effective architecture contains K/V projections and K normalization in the Prefiller; the Decoder has independent residual, FFN/MoE, and query-side parameters but does not produce an additional cache required by later positions.

The Decoder is initialized through a parameter-free RMS-normalized bridge:

DD4

with

DD5

The final Decoder state is passed through the shared output normalization and untied output head. The bridge and final-state RMS operations have no trainable parameters, accumulate in FP32, and return the input dtype in the reported implementation.

3. Execution across training, prefilling, and decoding

Training

After expansion, both towers process every supervised position. The Prefiller processes the sequence and generates layer-wise DD6 tensors. The Decoder computes predictions at all training positions, and gradients flow through its attention reads back into the Prefiller through the reused KV tensors. The Prefiller is therefore jointly trained after expansion rather than frozen.

For source and continuation token counts DD7, and forward FLOPs per token DD8, the paper estimates training cost as

DD9

where the factor 3 approximates forward plus backward computation. The theoretical accounting excludes optimizer cost, communication, checkpoint I/O, rematerialization, conversion overhead, normalization and residual micro-operations, and auxiliary loss terms.

Bulk prompt prefilling

For a prompt (Mt,zt)=P(x1:t;θP),(M_t,z_t)=P(x_{1:t};\theta_P),0, the Prefiller processes all (Mt,zt)=P(x1:t;θP),(M_t,z_t)=P(x_{1:t};\theta_P),1 tokens and constructs the layer-wise cache

(Mt,zt)=P(x1:t;θP),(M_t,z_t)=P(x_{1:t};\theta_P),2

The Decoder is not evaluated at every prompt position. It is evaluated only at the final prompt position (Mt,zt)=P(x1:t;θP),(M_t,z_t)=P(x_{1:t};\theta_P),3, using the Prefiller KV cache to produce the first output-token distribution. Earlier Decoder computations can be omitted because the Decoder does not generate memory needed by later prompt positions. The bulk-prefill path is therefore source-sized apart from the final-position Decoder computation and shared output operations.

Autoregressive decoding

For each generated token, the Prefiller processes the token and appends its K/V representations to the cache. The Decoder then processes that token, reads the expanded KV cache, and produces the next-token logits. KITE therefore does not eliminate the additional Decoder computation during generation. Its advantage depends on amortizing this decode overhead over sufficiently substantial prompt-prefill work.

4. Upcycling and parameter construction

The evaluated SST model was expanded from a 33.819B-parameter source model. The construction procedure retained the source Transformer as the SST Prefiller, copied each source layer into the corresponding Decoder layer, shared the embedding, output normalization, and output head, and then continued training both towers jointly.

This is aligned source-to-target expansion rather than strict function-preserving Net2Net expansion. Adding the Decoder changes the prediction graph, so the expanded model experiences a loss increase at conversion before adapting during continuation training. The Prefiller and shared-parameter optimizer states are retained, while fresh optimizer states are initialized for the Decoder. Training continues on the inherited cosine learning-rate schedule without a separate Decoder warmup.

The source stage used 167.98B tokens and the continuation stage used 222.55B tokens, for 390.54B total tokens at the evaluated SST endpoint. Training used BF16, Muon and Adam parameter groups, 128 NVIDIA H800 GPUs, expert parallelism of 8, sequence length 4096, and a global batch size of 3072 sequences. The continuation stage comprised 17,687 updates after 13,350 source updates.

Model Total parameters Layers Decode-active body parameters per token
Source 33.819B 18 1.120B
67B SST 66.959B (Mt,zt)=P(x1:t;θP),(M_t,z_t)=P(x_{1:t};\theta_P),4 2.155B
47B classic 46.727B 20 1.477B
63B classic 62.691B 22 2.016B

The SST model contains 66.365B body parameters and 0.594B embedding-plus-output-head parameters. Its bulk-prefill active body count is 1.120B, corresponding to the Prefiller-only path. The active counts include selected Top-8 experts, shared experts, routers, and other body weights required for one token, while excluding the embedding and output head.

5. Training quality and scaling results

The 47B classic model defines the normalized theoretical training budget:

(Mt,zt)=P(x1:t;θP),(M_t,z_t)=P(x_{1:t};\theta_P),5

At their reported final endpoints, the 47B classic, 63B classic, and 67B SST models each use approximately the same relative theoretical training FLOPs. Their training-token counts are 443.16B, 334.62B, and 390.54B, respectively.

The final EMA-200 training losses were:

Model Training tokens Final training loss
47B classic 443.16B 1.6006
63B classic 334.62B 1.5921
67B SST 390.54B 1.5900

SST’s final loss was 0.0106 lower than the 47B baseline and 0.0021 lower than the 63B baseline at comparable cumulative theoretical training compute. Its loss increased at the expansion point and subsequently decreased during continuation training.

The reported downstream results were:

Task 47B classic 63B classic 67B SST
OpenBookQA 68.50 71.00 75.00
MMLU 58.44 59.39 60.54
GSM8K 56.56 55.04 58.38
MATH 33.06 33.14 34.36
HumanEval, few-shot 35.98 32.93 37.80
MBPP, 3-shot 50.40 50.80 51.40
BBH 50.36 52.56 52.74
ARXIV NLL 1.4521 1.4424 1.4421

The comparisons used the same tokenizer, data recipe, training protocol, and evaluation checkpoints. The reported results support the claim that the expanded two-tower model can improve quality at comparable cumulative theoretical training compute, although the experiments do not establish a general scaling law or isolate the causal contribution of each architectural component.

6. Inference-cost accounting, workload dependence, and limitations

KITE’s inference-cost advantage is workload-dependent. The paper uses an analytical active-parameter proxy rather than measured latency, throughput, energy, or serving cost. Let (Mt,zt)=P(x1:t;θP),(M_t,z_t)=P(x_{1:t};\theta_P),6 denote bulk-prefill active-body parameters normalized to the 47B classic model, (Mt,zt)=P(x1:t;θP),(M_t,z_t)=P(x_{1:t};\theta_P),7 denote decode-active parameters under the same normalization, and (Mt,zt)=P(x1:t;θP),(M_t,z_t)=P(x_{1:t};\theta_P),8 denote the prefill cost weight:

(Mt,zt)=P(x1:t;θP),(M_t,z_t)=P(x_{1:t};\theta_P),9

For the illustrative 75:25 prefill-to-decode mixture,

ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),0

The normalized proxy is 1 for the 47B classic model, 0.933 for SST, and 1.365 for the 63B classic model. SST is consequently estimated to be 6.7% cheaper than the 47B model and 31.6% cheaper than the 63B model under this proxy.

The reported break-even points are approximately:

  • Against 63B classic: 13.4% prefill and 86.6% decode.
  • Against 47B classic: 65.5% prefill and 34.5% decode.

Thus SST is estimated to be cheaper than the 47B model when prefill constitutes more than roughly 65.5% of weighted cost, and cheaper than the 63B model for much less input-heavy mixtures. Applying selected OpenRouter traffic charge shares of 67.4%–76.7% as proxy prefill weights gives estimated SST savings of 1.3%–7.9% versus the 47B model and 27.7%–32.5% versus the 63B model. These traffic statistics are price-based workload indicators rather than direct measurements of GPU prefill work, and the selected traffic is not labeled specifically as agentic usage.

KITE is particularly suited to workloads involving repeated long-context processing, including repeated agent turns, tool-use loops, accumulated conversation state, large documents, and large code contexts. In short-prompt or decode-dominated workloads, the additional Decoder computation may eliminate the advantage.

The architecture does not eliminate the Prefiller’s KV cache, reduce its number of cached tokens, or directly replace MQA, GQA, cross-layer attention, or KV-compression methods. The paper does not provide a complete byte-level KV-memory comparison, measured serving speedups, or measured maximum-context improvements. Running two towers may also introduce synchronization, memory-bandwidth, expert-routing, and parallelization overheads not captured by the active-parameter proxy.

A further limitation is component attribution. SST combines upcycling, two-tower execution, MoE routing, layer-wise KV reuse, and the specified attention pattern. The experiments do not determine how much of the observed improvement arises specifically from KV invariance, initialization, the source model, or other architectural choices. The results therefore establish an empirical instance of KITE rather than a complete theory of Transformer expansion.

KITE’s principal distinction from KV-only self-attention is architectural. Key-Value Transformer variants remove the independent query projection from self-attention and replace ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),1 with ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),2, yielding a lower-cost attention mechanism with symmetric pre-softmax affinities (Borji, 2023). KITE does not remove queries from the Decoder or alter the self-attention similarity rule in that manner. Instead, it preserves a source-sized KV-producing path and adds a separate KV-reading capacity-expansion path. Its central claim is consequently about where additional Transformer computation is placed, not about replacing QKV attention within an individual attention block.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to KV-Invariant Transformer Expansion (KITE).