Papers
Topics
Authors
Recent
Search
2000 character limit reached

Step Scale Transformer (SST)

Updated 30 September 2026
  • Step Scale Transformer (SST) is a two-tower decoder architecture designed to enhance large language models by separating prompt-prefill and generation costs, improving efficiency, and allowing more scalable usage in long context scenarios using a two-tower architecture.
  • Spacing the large language models with an additional decoder tower that allows it to efficiently use the previously created cache for the sequence thus taking advantage of the potential in-memory large scale processing.
  • Maintaining the parameters and memory use so that the overall memory usage case and cpu usage reduces in such a way that it retains the efficiency.

Step Scale Transformer (SST) is a two-tower decoder architecture for scaling agentic LLMs while reducing prompt-prefill computation. Introduced as a concrete implementation of KV-Invariant Transformer Expansion (KITE), SST expands a smaller trained Transformer into a larger model by adding a Decoder tower that reads, but does not produce, the historical key–value (KV) memory used by subsequent positions. The Prefiller tower processes the complete prompt and constructs the layer-wise KV cache, whereas the Decoder supplies additional predictive capacity at the prompt boundary and during autoregressive generation. The architecture is designed for workloads in which long-context prompt processing constitutes a substantial fraction of inference cost (Hu et al., 23 Sep 2026).

1. Architectural motivation and terminology

SST addresses three distinct computational costs in autoregressive language modeling: training, prompt prefilling, and token-by-token decoding. Conventional Transformers use the entire model to process every prompt token and to construct the KV cache. Sparse mixture-of-experts (MoE) layers reduce active computation relative to total parameter capacity, but the same active model remains involved in both prefill and decoding. Consequently, increasing model capacity can increase prompt-processing cost even when the output sequence is short.

This distinction is particularly relevant to agentic workloads, which may repeatedly process accumulated conversation histories, tool outputs, retrieved documents, observations, and intermediate reasoning traces. Selected OpenRouter traffic cited in the SST study assigns approximately 67.4–76.7% of estimated uncached-input-plus-output charges to uncached input. These figures are used to motivate input-heavy workload analysis, although the traffic is not separately labeled as agentic usage (Hu et al., 23 Sep 2026).

The underlying design principle is KV-Invariant Transformer Expansion (KITE). Here, “KV-invariant” refers to an architectural invariant rather than numerical preservation of KV parameters or values. The KV-producing path remains the smaller source-sized path, while newly added capacity is placed in a path that reads the existing memory but does not generate KV required by future positions. Both towers can nevertheless be trained jointly after expansion.

KITE can be expressed abstractly as a Prefiller–Decoder decomposition:

(Mt,zt)=P(x1:t;θP),(M_t,z_t)=P(x_{1:t};\theta_P),

ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),

where PP is the Prefiller, DD is the Decoder, MtM_t is reusable attention memory, ztz_t denotes optional Prefiller features, and ℓt\ell_t is the next-token logit vector. SST is the reported two-tower realization of this decomposition.

2. Two-tower SST architecture

SST contains two sequential Transformer towers of equal depth:

  1. The Prefiller tower processes the complete sequence and produces layer-wise keys and values.
  2. The Decoder tower reads the Prefiller’s KV tensors and provides additional token-local computation and predictive capacity.

In the reported configuration, each tower contains 18 Transformer layers, giving 36 Transformer blocks in total. Both towers use hidden width 2304, but their Transformer-block parameters are independent. SST is expanded from a 33.819-billion-parameter source model to a 66.959-billion-parameter model.

Model Layers Hidden width MoE layers Total parameters Active body parameters per decode token
Source 18 2304 16 33.819B 1.120B
SST $18+18$ 2304 $16+16$ 66.959B 2.155B
47B classic 20 2560 18 46.727B 1.477B
63B classic 22 2816 20 62.691B 2.016B

All listed models use 512 routed experts per MoE layer, Top-8 expert selection, SSSF attention, head dimension 128, eight KV heads, full-attention RoPE dimension 64, and sliding-attention RoPE dimension 128. SSSF attention is described as three sliding-window layers followed by one full-attention layer, with the remaining pattern determined by depth. The sliding window is 512 tokens in the training FLOP accounting.

For Decoder layer ll and token position ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),0, the query is generated from the Decoder hidden state:

ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),1

The Decoder then attends to the corresponding Prefiller-layer KV:

ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),2

Here, ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),3 and ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),4 are generated by Prefiller layer ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),5, ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),6 is the Decoder query projection, and ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),7 specifies the positions allowed by the causal or sliding attention mask. The layer alignment is essential: Decoder layer ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),8 reads KV from Prefiller layer ℓt=D(xt,zt,Mt;θD),\ell_t=D(x_t,z_t,M_t;\theta_D),9, rather than from a single globally shared KV representation.

The Decoder receives the shared token embedding and the final Prefiller state through a parameter-free RMS-normalized bridge:

PP0

where

PP1

The two RMS operations have no trainable parameters and accumulate in FP32. SST shares one token embedding, one output normalization, and one output head between the towers; the embedding and output head are untied. Its output is

PP2

Gradients reach the Prefiller through two routes: the bridge involving the final Prefiller state and the Decoder attention path involving the reused Prefiller KV.

3. Information flow during training and inference

During training, the Prefiller processes every token and produces all layer-wise KV tensors. The Decoder also computes every supervised position and reads the corresponding Prefiller KV. Both towers therefore participate in full-sequence training, and gradients propagate through the Decoder, the KV-producing Prefiller path, and shared parameters.

Prompt prefilling uses a different execution pattern. For a prompt PP3, the Prefiller processes all PP4 tokens and constructs the layer-wise KV cache. The Decoder is evaluated only at the final prompt position PP5, because that position is required to predict the first generated token. Earlier Decoder positions can be omitted because Decoder operations other than attention to Prefiller KV are token-local.

The omitted historical Decoder states do not affect the Decoder state at the final prompt position. That state depends on the token embedding at position PP6, the final Prefiller state at position PP7, and the Prefiller-produced KV memory permitted by the attention mask. This dependency structure is the basis for reducing bulk-prefill computation.

Autoregressive decoding requires both towers for each newly generated token:

  1. The new token passes through the Prefiller.
  2. The Prefiller extends the KV cache.
  3. The Decoder processes the new token.
  4. The Decoder reads the relevant Prefiller KV and generates next-token logits.

Thus, the statement that KV caching “only depends on the smaller tower” has a specific meaning. The KV cache is produced exclusively by the Prefiller; the Decoder does not add a second historical KV stream. However, the Decoder still runs at the prompt boundary and on every generated token. SST reduces bulk-prefill computation but does not reduce decoding computation relative to the source-sized model.

The architecture imposes several dependency constraints:

  • Decoder-retained positions must not depend on omitted historical Decoder positions.
  • Future memory must be generated entirely by the Prefiller.
  • Decoder access to historical information must occur through aligned Prefiller KV.
  • The Prefiller–Decoder interface must remain compatible after expansion.
  • Position IDs and attention masks must remain consistent with the source model.

If future memory depended on historical Decoder states, the earlier Decoder computations could not be omitted during prefill.

4. Upcycling and training procedure

SST is constructed through a two-stage expansion procedure. First, an 18-layer source model is trained. The reported source model contains 33.819 billion parameters and is trained for 13,350 updates over 167.98 billion tokens.

At conversion, every source layer is copied into the corresponding Prefiller layer and also into the corresponding Decoder layer. The embedding, output normalization, and output head are shared. Prefiller and shared-parameter optimizer states are retained, whereas Decoder optimizer states are initialized freshly.

This aligned initialization is not function-preserving. Because the Decoder changes the prediction graph, the expanded model experiences a loss increase at conversion before adapting during continued training.

The two towers are then trained jointly for 17,687 continuation updates over 222.55 billion continuation tokens. Including the source stage, the reported total is 390.54 billion tokens. The continuation uses BF16 parameters, Muon and Adam parameter groups, 128 NVIDIA H800 GPUs, expert parallelism of 8, sequence length 4096, and global batch size 3,072 sequences.

The source cosine learning-rate schedule is continued rather than restarted with a separate Decoder warmup. The source peak base learning rate is approximately

PP8

with continuation anchor and minimum values of approximately

PP9

For continuation progress DD0 and configured continuation horizon DD1, the schedule is

DD2

DD3

The evaluated source model and SST use the same tokenizer, data recipe, broad training protocol, and evaluation setup as the classic baselines. The supplied study does not establish whether aligned copying is superior to random Decoder initialization, frozen-Prefiller training, independently initialized towers, separate Decoder warmup, or distillation-based expansion, because those alternatives are not reported as systematic ablations.

5. Training and inference cost

For source and continuation token counts DD4 and DD5, and forward FLOPs per token DD6 and DD7, the reported training-compute estimate is

DD8

where the factor of 3 approximates forward plus backward computation.

The 47B classic model defines the reference training budget at approximately

DD9

Model Training tokens Relative theoretical training FLOPs
47B classic 443.16B 1.000
63B classic 334.62B Approximately 1.000
67B SST 390.54B Approximately 1.000

Reported forward FLOPs per token are approximately 3.087252 billion for the source model, 3.887488 billion for the 47B classic model, 5.148513 billion for the 63B classic model, and 5.410678 billion for SST. The estimator uses sequence length 4096, sliding window 512, and padded vocabulary 128,896. It includes projections, attention products, embedding, output head, MoE and FFN computation, routers, and activation terms, but excludes normalization, RoPE, residuals, optimizer cost, communication, checkpointing, and other overheads.

SST contains 66.959 billion total parameters, including 66.365 billion body parameters and 0.594 billion embedding-plus-output-head parameters. Its active body-parameter count is 2.155 billion per decoding token, of which 1.120 billion belongs to the Prefiller-only path.

Inference cost is evaluated through an analytical proxy rather than measured latency or throughput. Let MtM_t0 denote bulk-prefill active-body cost normalized to the 47B classic model, MtM_t1 denote decode active-body cost under the same normalization, and MtM_t2 denote the prefill fraction. The proxy is

MtM_t3

For the illustrative workload mix MtM_t4, the reported normalized values are:

Model Bulk-prefill proxy Decode proxy MtM_t5
47B classic 1.000 1.000 1.000
SST Approximately MtM_t6 MtM_t7 0.933
63B classic Approximately 1.000 relative to its full-body ratio MtM_t8 1.365

Under this proxy, SST is estimated to be 6.7% lower than the 47B classic model and 31.6% lower than the 63B classic model. The reported break-even workload mixtures are approximately 65.5:34.5 prefill:decode for SST versus the 47B classic model and 13.4:86.6 prefill:decode for SST versus the 63B classic model.

The estimates are not measured end-to-end serving improvements. They do not model GPU utilization, memory bandwidth, kernel fusion, communication, KV-cache movement, expert load imbalance, scheduling, batching, interconnect topology, or prompt-boundary Decoder overhead. Actual performance therefore depends on runtime implementation and hardware execution.

6. Empirical results, comparisons, and limitations

At the final reported checkpoints, SST achieves lower EMA-200 training loss than both classic baselines at comparable theoretical training compute.

Model Tokens EMA-200 training loss
47B classic 443.16B 1.6006
63B classic 334.62B 1.5921
67B SST 390.54B 1.5900

SST’s final loss is 0.0106 lower than the 47B classic model and 0.0021 lower than the 63B classic model. Its loss rises at conversion and subsequently decreases during joint continuation.

On the reported downstream tasks, SST scores higher than both classic baselines on all seven listed task metrics and has the lowest reported ARXIV NLL.

Task 47B classic 63B classic 67B SST
OpenBookQA 68.50 71.00 75.00
MMLU 58.44 59.39 60.54
GSM8K 56.56 55.04 58.38
MATH 33.06 33.14 34.36
HumanEval 35.98 32.93 37.80
MBPP 50.40 50.80 51.40
BBH 50.36 52.56 52.74
ARXIV NLL 1.4521 1.4424 1.4421

SST is related to several other efficiency strategies but has a distinct objective. MoE increases total capacity while activating only a subset of experts; SST is orthogonal to MoE and incorporates MoE layers in both towers. Progressive upcycling reduces training cost by starting from a smaller model, whereas KITE additionally preserves a smaller bulk-prefill path. YOCO-style designs separate cache-producing and cache-consuming computation, but SST combines this execution separation with source-model expansion, aligned layer-wise KV reuse, and joint post-expansion training. MQA, GQA, and cross-layer attention reduce KV storage or share KV across heads or layers; KITE instead targets the computation required to construct KV and is potentially complementary to those methods.

The principal architectural limitation is the dependency separation required for prefill savings. SST constrains information flow by ensuring that historical Decoder states are unnecessary for future memory construction. Its decode cost is higher than that of the 47B classic model because every generated token uses both towers. Decode-dominated workloads can therefore erase the prefill advantage.

Other limitations concern evaluation and systems realization. The reported inference savings are active-parameter proxies rather than measured latency, throughput, energy, or cost per request. The supplied results do not establish optimal Prefiller–Decoder depth ratios, hidden widths, KV dimensions, KV-head counts, or alternative initialization procedures. The KITE benefit may also require specialized runtime support for separate Prefiller and Decoder execution, layer-aligned KV routing, prompt-boundary-only Decoder evaluation during prefill, and potentially distinct parallelization strategies for prefill and decoding.

Within the reported evidence, SST demonstrates a specific scaling strategy: increase predictive capacity after the KV-producing path while preserving a smaller model for bulk prompt processing. Its measured training and benchmark results support the reported SST configuration, while the broader claim that KITE constitutes a generally superior scaling paradigm remains dependent on larger models, systematic architectural ablations, measured multi-GPU serving experiments, and evaluations under realistic long-context agentic workloads.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Step Scale Transformer (SST).