---
title: Step Scale Transformer (SST)
url: https://www.emergentmind.com/topics/step-scale-transformer-sst
type: topic
---

# Step Scale Transformer (SST)

Step Scale Transformer (SST) is a two-tower decoder architecture for scaling agentic large language models while reducing prompt-prefill computation. Introduced as a concrete implementation of KV-Invariant Transformer Expansion (KITE), SST expands a smaller trained Transformer into a larger model by adding a Decoder tower that reads, but does not produce, the historical key–value (KV) memory used by subsequent positions. The Prefiller tower processes the complete prompt and constructs the layer-wise KV cache, whereas the Decoder supplies additional predictive capacity at the prompt boundary and during autoregressive generation. The architecture is designed for workloads in which long-context prompt processing constitutes a substantial fraction of inference cost [2609.27294].

## 1. Architectural motivation and terminology

SST addresses three distinct computational costs in autoregressive language modeling: training, prompt prefilling, and token-by-token decoding. Conventional Transformers use the entire model to process every prompt token and to construct the KV cache. Sparse mixture-of-experts (MoE) layers reduce active computation relative to total parameter capacity, but the same active model remains involved in both prefill and decoding. Consequently, increasing model capacity can increase prompt-processing cost even when the output sequence is short.

This distinction is particularly relevant to agentic workloads, which may repeatedly process accumulated conversation histories, tool outputs, retrieved documents, observations, and intermediate reasoning traces. Selected OpenRouter traffic cited in the SST study assigns approximately 67.4–76.7% of estimated uncached-input-plus-output charges to uncached input. These figures are used to motivate input-heavy workload analysis, although the traffic is not separately labeled as agentic usage [2609.27294].

The underlying design principle is **KV-Invariant Transformer Expansion (KITE)**. Here, “KV-invariant” refers to an architectural invariant rather than numerical preservation of KV parameters or values. The KV-producing path remains the smaller source-sized path, while newly added capacity is placed in a path that reads the existing memory but does not generate KV required by future positions. Both towers can nevertheless be trained jointly after expansion.

KITE can be expressed abstractly as a Prefiller–Decoder decomposition:

$$
(M_t,z_t)=P(x_{1:t};\theta_P),
$$

$$
\ell_t=D(x_t,z_t,M_t;\theta_D),
$$

where $P$ is the Prefiller, $D$ is the Decoder, $M_t$ is reusable attention memory, $z_t$ denotes optional Prefiller features, and $\ell_t$ is the next-token logit vector. SST is the reported two-tower realization of this decomposition.

## 2. Two-tower SST architecture

SST contains two sequential Transformer towers of equal depth:

1. The **Prefiller tower** processes the complete sequence and produces layer-wise keys and values.
2. The **Decoder tower** reads the Prefiller’s KV tensors and provides additional token-local computation and predictive capacity.

In the reported configuration, each tower contains 18 Transformer layers, giving 36 Transformer blocks in total. Both towers use hidden width 2304, but their Transformer-block parameters are independent. SST is expanded from a 33.819-billion-parameter source model to a 66.959-billion-parameter model.

| Model | Layers | Hidden width | MoE layers | Total parameters | Active body parameters per decode token |
|---|---:|---:|---:|---:|---:|
| Source | 18 | 2304 | 16 | 33.819B | 1.120B |
| SST | \(18+18\) | 2304 | \(16+16\) | 66.959B | 2.155B |
| 47B classic | 20 | 2560 | 18 | 46.727B | 1.477B |
| 63B classic | 22 | 2816 | 20 | 62.691B | 2.016B |

All listed models use 512 routed experts per MoE layer, Top-8 expert selection, SSSF attention, head dimension 128, eight KV heads, full-attention RoPE dimension 64, and sliding-attention RoPE dimension 128. SSSF attention is described as three sliding-window layers followed by one full-attention layer, with the remaining pattern determined by depth. The sliding window is 512 tokens in the training FLOP accounting.

For Decoder layer $l$ and token position $t$, the query is generated from the Decoder hidden state:

$$
q^{D}_{l,t}=Q_l^{D}(h^{D}_{l-1,t}).
$$

The Decoder then attends to the corresponding Prefiller-layer KV:

$$
a^{D}_{l,t}
=
\operatorname{Attn}
\left(
q^{D}_{l,t},
K_l^{P}[\mathcal I_l(t)],
V_l^{P}[\mathcal I_l(t)]
\right).
$$

Here, $K_l^P$ and $V_l^P$ are generated by Prefiller layer $l$, $Q_l^D$ is the Decoder query projection, and $\mathcal I_l(t)$ specifies the positions allowed by the causal or sliding attention mask. The layer alignment is essential: Decoder layer $l$ reads KV from Prefiller layer $l$, rather than from a single globally shared KV representation.

The Decoder receives the shared token embedding and the final Prefiller state through a parameter-free RMS-normalized bridge:

$$
h^{D}_{0,t}
=
\operatorname{RMS}_{\epsilon}(e_t)
+
\operatorname{RMS}_{\epsilon}(h^{P}_{L,t}),
$$

where

$$
\operatorname{RMS}_{\epsilon}(v)
=
\frac{v}{\sqrt{\operatorname{mean}(v^2)+\epsilon}},
\qquad
\epsilon=10^{-5}.
$$

The two RMS operations have no trainable parameters and accumulate in FP32. SST shares one token embedding, one output normalization, and one output head between the towers; the embedding and output head are untied. Its output is

$$
\ell_t
=
W_{\mathrm{out}}\operatorname{Norm}_{\mathrm{out}}
\left(h^{D}_{L,t}\right).
$$

Gradients reach the Prefiller through two routes: the bridge involving the final Prefiller state and the Decoder attention path involving the reused Prefiller KV.

## 3. Information flow during training and inference

During training, the Prefiller processes every token and produces all layer-wise KV tensors. The Decoder also computes every supervised position and reads the corresponding Prefiller KV. Both towers therefore participate in full-sequence training, and gradients propagate through the Decoder, the KV-producing Prefiller path, and shared parameters.

Prompt prefilling uses a different execution pattern. For a prompt $x_{1:n}$, the Prefiller processes all $n$ tokens and constructs the layer-wise KV cache. The Decoder is evaluated only at the final prompt position $n$, because that position is required to predict the first generated token. Earlier Decoder positions can be omitted because Decoder operations other than attention to Prefiller KV are token-local.

The omitted historical Decoder states do not affect the Decoder state at the final prompt position. That state depends on the token embedding at position $n$, the final Prefiller state at position $n$, and the Prefiller-produced KV memory permitted by the attention mask. This dependency structure is the basis for reducing bulk-prefill computation.

Autoregressive decoding requires both towers for each newly generated token:

1. The new token passes through the Prefiller.
2. The Prefiller extends the KV cache.
3. The Decoder processes the new token.
4. The Decoder reads the relevant Prefiller KV and generates next-token logits.

Thus, the statement that KV caching “only depends on the smaller tower” has a specific meaning. The KV cache is produced exclusively by the Prefiller; the Decoder does not add a second historical KV stream. However, the Decoder still runs at the prompt boundary and on every generated token. SST reduces bulk-prefill computation but does not reduce decoding computation relative to the source-sized model.

The architecture imposes several dependency constraints:

- Decoder-retained positions must not depend on omitted historical Decoder positions.
- Future memory must be generated entirely by the Prefiller.
- Decoder access to historical information must occur through aligned Prefiller KV.
- The Prefiller–Decoder interface must remain compatible after expansion.
- Position IDs and attention masks must remain consistent with the source model.

If future memory depended on historical Decoder states, the earlier Decoder computations could not be omitted during prefill.

## 4. Upcycling and training procedure

SST is constructed through a two-stage expansion procedure. First, an 18-layer source model is trained. The reported source model contains 33.819 billion parameters and is trained for 13,350 updates over 167.98 billion tokens.

At conversion, every source layer is copied into the corresponding Prefiller layer and also into the corresponding Decoder layer. The embedding, output normalization, and output head are shared. Prefiller and shared-parameter optimizer states are retained, whereas Decoder optimizer states are initialized freshly.

This aligned initialization is not function-preserving. Because the Decoder changes the prediction graph, the expanded model experiences a loss increase at conversion before adapting during continued training.

The two towers are then trained jointly for 17,687 continuation updates over 222.55 billion continuation tokens. Including the source stage, the reported total is 390.54 billion tokens. The continuation uses BF16 parameters, Muon and Adam parameter groups, 128 NVIDIA H800 GPUs, expert parallelism of 8, sequence length 4096, and global batch size 3,072 sequences.

The source cosine learning-rate schedule is continued rather than restarted with a separate Decoder warmup. The source peak base learning rate is approximately

$$
1.03302\times 10^{-3},
$$

with continuation anchor and minimum values of approximately

$$
\eta_{\mathrm{anchor}}=5.96548\times10^{-4},
\qquad
\eta_{\min}=1.03302\times10^{-4}.
$$

For continuation progress $u$ and configured continuation horizon $U$, the schedule is

$$
q(u)=q_0+(1-q_0)\min(u/U,1),
$$

$$
\eta(u)=
\eta_{\min}
+
(\eta_{\mathrm{anchor}}-\eta_{\min})
\frac{1+\cos(\pi q(u))}
{1+\cos(\pi q_0)}.
$$

The evaluated source model and SST use the same tokenizer, data recipe, broad training protocol, and evaluation setup as the classic baselines. The supplied study does not establish whether aligned copying is superior to random Decoder initialization, frozen-Prefiller training, independently initialized towers, separate Decoder warmup, or distillation-based expansion, because those alternatives are not reported as systematic ablations.

## 5. Training and inference cost

For source and continuation token counts $D_s$ and $D_c$, and forward FLOPs per token $f_s$ and $f_c$, the reported training-compute estimate is

$$
C_{\mathrm{train}}
=
3(D_s f_s+D_c f_c),
$$

where the factor of 3 approximates forward plus backward computation.

The 47B classic model defines the reference training budget at approximately

$$
5.16836\times 10^{21}\ \text{FLOPs}.
$$

| Model | Training tokens | Relative theoretical training FLOPs |
|---|---:|---:|
| 47B classic | 443.16B | 1.000 |
| 63B classic | 334.62B | Approximately 1.000 |
| 67B SST | 390.54B | Approximately 1.000 |

Reported forward FLOPs per token are approximately 3.087252 billion for the source model, 3.887488 billion for the 47B classic model, 5.148513 billion for the 63B classic model, and 5.410678 billion for SST. The estimator uses sequence length 4096, sliding window 512, and padded vocabulary 128,896. It includes projections, attention products, embedding, output head, MoE and FFN computation, routers, and activation terms, but excludes normalization, RoPE, residuals, optimizer cost, communication, checkpointing, and other overheads.

SST contains 66.959 billion total parameters, including 66.365 billion body parameters and 0.594 billion embedding-plus-output-head parameters. Its active body-parameter count is 2.155 billion per decoding token, of which 1.120 billion belongs to the Prefiller-only path.

Inference cost is evaluated through an analytical proxy rather than measured latency or throughput. Let $p$ denote bulk-prefill active-body cost normalized to the 47B classic model, $d$ denote decode active-body cost under the same normalization, and $w$ denote the prefill fraction. The proxy is

$$
C_{\mathrm{infer}}(w)=wp+(1-w)d.
$$

For the illustrative workload mix $w=0.75$, the reported normalized values are:

| Model | Bulk-prefill proxy | Decode proxy | \(0.75p+0.25d\) |
|---|---:|---:|---:|
| 47B classic | 1.000 | 1.000 | 1.000 |
| SST | Approximately \(1.120/1.477\) | \(2.155/1.477\) | 0.933 |
| 63B classic | Approximately 1.000 relative to its full-body ratio | \(2.016/1.477\) | 1.365 |

Under this proxy, SST is estimated to be 6.7% lower than the 47B classic model and 31.6% lower than the 63B classic model. The reported break-even workload mixtures are approximately 65.5:34.5 prefill:decode for SST versus the 47B classic model and 13.4:86.6 prefill:decode for SST versus the 63B classic model.

The estimates are not measured end-to-end serving improvements. They do not model GPU utilization, memory bandwidth, kernel fusion, communication, KV-cache movement, expert load imbalance, scheduling, batching, interconnect topology, or prompt-boundary Decoder overhead. Actual performance therefore depends on runtime implementation and hardware execution.

## 6. Empirical results, comparisons, and limitations

At the final reported checkpoints, SST achieves lower EMA-200 training loss than both classic baselines at comparable theoretical training compute.

| Model | Tokens | EMA-200 training loss |
|---|---:|---:|
| 47B classic | 443.16B | 1.6006 |
| 63B classic | 334.62B | 1.5921 |
| 67B SST | 390.54B | **1.5900** |

SST’s final loss is 0.0106 lower than the 47B classic model and 0.0021 lower than the 63B classic model. Its loss rises at conversion and subsequently decreases during joint continuation.

On the reported downstream tasks, SST scores higher than both classic baselines on all seven listed task metrics and has the lowest reported ARXIV NLL.

| Task | 47B classic | 63B classic | 67B SST |
|---|---:|---:|---:|
| OpenBookQA | 68.50 | 71.00 | **75.00** |
| MMLU | 58.44 | 59.39 | **60.54** |
| GSM8K | 56.56 | 55.04 | **58.38** |
| MATH | 33.06 | 33.14 | **34.36** |
| HumanEval | 35.98 | 32.93 | **37.80** |
| MBPP | 50.40 | 50.80 | **51.40** |
| BBH | 50.36 | 52.56 | **52.74** |
| ARXIV NLL | 1.4521 | 1.4424 | **1.4421** |

SST is related to several other efficiency strategies but has a distinct objective. MoE increases total capacity while activating only a subset of experts; SST is orthogonal to MoE and incorporates MoE layers in both towers. Progressive upcycling reduces training cost by starting from a smaller model, whereas KITE additionally preserves a smaller bulk-prefill path. YOCO-style designs separate cache-producing and cache-consuming computation, but SST combines this execution separation with source-model expansion, aligned layer-wise KV reuse, and joint post-expansion training. MQA, GQA, and cross-layer attention reduce KV storage or share KV across heads or layers; KITE instead targets the computation required to construct KV and is potentially complementary to those methods.

The principal architectural limitation is the dependency separation required for prefill savings. SST constrains information flow by ensuring that historical Decoder states are unnecessary for future memory construction. Its decode cost is higher than that of the 47B classic model because every generated token uses both towers. Decode-dominated workloads can therefore erase the prefill advantage.

Other limitations concern evaluation and systems realization. The reported inference savings are active-parameter proxies rather than measured latency, throughput, energy, or cost per request. The supplied results do not establish optimal Prefiller–Decoder depth ratios, hidden widths, KV dimensions, KV-head counts, or alternative initialization procedures. The KITE benefit may also require specialized runtime support for separate Prefiller and Decoder execution, layer-aligned KV routing, prompt-boundary-only Decoder evaluation during prefill, and potentially distinct parallelization strategies for prefill and decoding.

Within the reported evidence, SST demonstrates a specific scaling strategy: increase predictive capacity after the KV-producing path while preserving a smaller model for bulk prompt processing. Its measured training and benchmark results support the reported SST configuration, while the broader claim that KITE constitutes a generally superior scaling paradigm remains dependent on larger models, systematic architectural ablations, measured multi-GPU serving experiments, and evaluations under realistic long-context agentic workloads.

Source: https://www.emergentmind.com/topics/step-scale-transformer-sst