---
title: 'NCP-ArchPreview: Latent-Space Language Model'
url: https://www.emergentmind.com/topics/ncp-archpreview
type: topic
---

# NCP-ArchPreview: Latent-Space Language Model

NCP-ArchPreview is a latent-space language-model architecture introduced to extend causal autoregressive pretraining beyond standard next-token prediction (NTP). It jointly predicts surface tokens and discrete, product-quantized concepts spanning multiple tokens. The architecture comprises a token encoder, a compressed concept module, and a token decoder; predicted concepts are causally fed back into token-level generation. A technical report describes an 8.9-billion-parameter implementation trained on 5.73 trillion Dolma-3 tokens, together with experiments on pretraining efficiency, downstream performance, lightweight domain adaptation, and speculative decoding [2609.10262].

## 1. Motivation and conceptual basis

Standard causal language modeling optimizes the conditional probability of the next token,

$$
p(x_{t+1}\mid x_{\leq t}),
$$

through the NTP objective. NCP-ArchPreview retains this objective but adds Next Concept Prediction (NCP), in which the model predicts a discrete latent representation associated with a future span of tokens. The motivation is that hidden states in language models already encode abstractions extending across multiple token positions, including entities, propositions, syntactic structures, mathematical substeps, and other semantic regularities. NTP supervises these abstractions only indirectly through individual token likelihoods.

NCP-ArchPreview therefore introduces an explicit predictive hierarchy:

$$
\text{tokens}
\rightarrow
\text{continuous concepts}
\rightarrow
\text{discrete concept vocabulary}
\rightarrow
\text{predicted future concepts}
\rightarrow
\text{token generation}.
$$

The concept objective is not implemented as a replacement for token prediction. Instead, predicted concepts influence the token decoder through causal feedback, while the final output remains a standard autoregressive token sequence. This distinguishes NCP-ArchPreview from ordinary multi-token prediction: the targets are representations summarizing multiple token positions rather than collections of independent future tokens.

The architecture has three principal computational stages:

1. **Token Encoder**: produces token-resolution hidden states.
2. **Concept Module**: processes a sequence compressed by a factor of four and predicts future quantized concepts.
3. **Token Decoder**: returns to token resolution and uses both token-level states and predicted concepts to produce next-token distributions.

The reported implementation has approximately 8.94 billion parameters and a maximum context length of 8,192 tokens. Its principal dimensions include hidden size 4,096, FFN size 11,008, 32 attention heads, 32 key-value groups, head dimension 128, vocabulary size 100,278, SwiGLU activation, RMSNorm with $\epsilon=10^{-6}$, RoPE with base 500,000, a token-level local-attention window of 4,096, and zero dropout.

## 2. Network architecture and hierarchical residual routing

The token-level Transformer consists of 16 encoder layers and 16 decoder layers. Between them is an eight-layer causal Concept Module operating at one-quarter of the token sequence length. The concept module consequently adds approximately eight ordinary Transformer blocks in parameter cost while requiring substantially less full-resolution computation.

The model extends ordinary residual propagation with two mechanisms:

- **Intra-Module Residual Connections (IRC)** dynamically combine representations from different depths within a module.
- **Cross-Module Residual Connections (CRC)** transmit representations between the Token Encoder, Concept Module, and Token Decoder.

For module $s$, let $\mathbf H_\ell^s$ denote the input to layer $\ell$, and let $\mathbf R_\ell^s=F_\ell^s(\mathbf H_\ell^s)$ be its block output. Standard residual propagation is

$$
\mathbf H_{\ell+1}^s=\mathbf H_\ell^s+\mathbf R_\ell^s.
$$

IRC instead forms a set of representations from multiple depths,

$$
\mathcal X_\ell^s=
\left\{
\mathbf H_1^s,\,
\mathbf H_1^s+\mathbf R_1^s,\,
\ldots,\,
\mathbf H_\ell^s+\mathbf R_\ell^s
\right\},
$$

and computes mixing coefficients through an MLP:

$$
\mathbf w_\ell^s=\operatorname{MLP}_\ell^s(\mathbf R_\ell^s).
$$

The next state is

$$
\mathbf H_{\ell+1}^s=
\sum_{j=1}^{\ell+1}w_{\ell,j}^s\mathbf X_{\ell,j}^s.
$$

The coefficients are not normalized, so signed combinations are possible. Initialization sets the final coefficient to one and all other coefficients to zero, initially recovering standard residual-Transformer behavior.

CRC aligns source representations to the target sequence granularity through repetition or chunk pooling. For source states $\{\mathbf S_j^u\}_{j=1}^K$ and target state $\mathbf T_\ell^s$, it computes

$$
\boldsymbol{\alpha}_\ell^{s\leftarrow u}
=
\operatorname{softmax}
\left(
\operatorname{MLP}_\ell^{s\leftarrow u}(\mathbf T_\ell^s)
\right),
$$

followed by

$$
\mathbf M_\ell^{s\leftarrow u}
=
\sum_{j=1}^{K}
\alpha_{\ell,j}^{s\leftarrow u}
\operatorname{LN}(\mathbf S_j^u)
$$

and

$$
\mathbf T_{\ell,\mathrm{out}}^s
=
\mathbf T_\ell^s+
\mathbf D_\ell^{s\leftarrow u}
\odot
\mathbf M_\ell^{s\leftarrow u},
$$

where $\mathbf D_\ell^{s\leftarrow u}$ is a learned diagonal scaling vector. The principal CRC paths are Token Encoder $\rightarrow$ Concept Module, Token Encoder $\rightarrow$ Token Decoder, and Concept Module $\rightarrow$ Token Decoder. Cross-module scales are initialized small so that training begins close to the original OLMo-style model.

## 3. Concept formation and product-quantized vocabulary

For an input sequence $x_{1:T}$, the Token Encoder produces hidden states

$$
\mathbf h_{1:T}
=
\operatorname{TokenEncoder}_{\theta_e}(x_{1:T}),
\qquad
\mathbf h_t\in\mathbb R^d.
$$

These states are passed toward the decoder, compressed into concepts, and used to learn the concept codebooks.

The compression factor is $k=4$. Every contiguous group of four token states is mean-pooled:

$$
\mathbf c_m
=
\frac{1}{k}
\sum_{i=1}^{k}
\mathbf h_{(m-1)k+i},
\qquad
m=1,\ldots,M,
$$

where

$$
M=\left\lfloor\frac{T}{k}\right\rfloor.
$$

Each concept therefore summarizes a fixed four-token span. The reported model does not use input-adaptive span boundaries; concept granularity is fixed at four tokens.

Each continuous concept is divided into $S=32$ segments:

$$
\mathbf c_m
=
\operatorname{concat}
\left(
\mathbf c_m^1,\ldots,\mathbf c_m^S
\right),
\qquad
\mathbf c_m^s\in\mathbb R^{d/S}.
$$

With $d=4096$, each segment has dimension $128$. Every segment has an independent codebook,

$$
\mathcal E^s=
\{\mathbf e_1^s,\ldots,\mathbf e_N^s\},
\qquad
N=128.
$$

The codeword selected for segment $s$ of concept $m$ is

$$
n_m^s
=
\arg\min_n
\left\|
\mathbf c_m^s-\mathbf e_n^s
\right\|_2^2,
$$

with quantized segment

$$
\mathbf d_m^s=\mathbf e_{n_m^s}^s.
$$

The complete quantized concept is

$$
\mathbf d_m
=
\operatorname{concat}
\left(
\mathbf d_m^1,\ldots,\mathbf d_m^S
\right).
$$

Although each segment codebook has only 128 entries, product quantization provides a latent vocabulary with capacity

$$
N^S=128^{32}
$$

possible code combinations. The model does not instantiate this as one monolithic table. Instead, it predicts 32 categorical distributions and combines their expected codewords.

The VQ loss is

$$
\mathcal L_{\mathrm{VQ}}
=
\frac{1}{MS}
\sum_{m=1}^{M}
\sum_{s=1}^{S}
\left\|
\operatorname{sg}(\mathbf c_m^s)-\mathbf d_m^s
\right\|_2^2,
$$

where $\operatorname{sg}$ denotes stop-gradient. This moves selected codebook vectors toward encoder-generated concepts without directly propagating the VQ reconstruction gradient through the Token Encoder. The encoder can nevertheless receive gradients through NCP and NTP.

## 4. Next Concept Prediction and causal feedback

The Concept Module receives preceding continuous concepts,

$$
\mathbf c_{<m}
=
(\mathbf c_1,\ldots,\mathbf c_{m-1}),
$$

and produces a latent state

$$
\mathbf u_m
=
\operatorname{ConceptModule}_{\theta_c}(\mathbf c_{<m}).
$$

For each product-quantization segment $s$, a prediction head outputs a distribution over 128 codewords:

$$
\boldsymbol\pi_m^s
=
\operatorname{softmax}
\left(
\operatorname{PredictionHead}_c^s(\mathbf u_m)
\right).
$$

The predicted segment is the expected codeword,

$$
\widehat{\mathbf c}_m^s
=
\sum_{n=1}^{N}
\pi_{m,n}^s\mathbf e_n^s,
$$

and the complete predicted concept is

$$
\widehat{\mathbf c}_m
=
\operatorname{concat}
\left(
\widehat{\mathbf c}_m^1,\ldots,\widehat{\mathbf c}_m^S
\right).
$$

Expectation over codewords avoids nondifferentiable argmax selection and constrains predictions to combinations of learned codebook entries. The NCP target is the continuous pooled representation, detached from the gradient:

$$
\mathcal L_{\mathrm{NCP}}
=
\frac{1}{M-1}
\sum_{m=2}^{M}
\left\|
\widehat{\mathbf c}_m
-
\operatorname{sg}(\mathbf c_m)
\right\|_2^2.
$$

An equivalent segment-normalized form divides by $(M-1)S$ and sums segment-wise squared errors. Gradients flow through the Concept Module, prediction heads, codebook-dependent prediction pathway, and preceding concept history, but not through the target concept.

To influence token generation, predicted concepts are repeated four times and shifted forward to preserve causality. For token position $t$, the broadcast concept signal is

$$
\mathbf b_t=
\begin{cases}
\mathbf 0, & 1\leq t<k,\\
\widehat{\mathbf c}_{\lfloor(t-k)/k\rfloor+2},
& k\leq t<T.
\end{cases}
$$

The zero prefix prevents the decoder from using a concept derived from tokens whose prediction it is supervising. The token representation supplied to the decoder is

$$
\widetilde{\mathbf h}_t
=
\mathbf h_t+\mathbf b_t.
$$

The Token Decoder then models

$$
p(x_{t+1}\mid x_{\leq t})
=
P_{\theta_d}
\left(
\widetilde{\mathbf h}_{\leq t}
\right).
$$

The Concept Module-to-Decoder CRC path uses the same causal shift. Consequently, both direct concept injection and cross-module concept residuals are designed to avoid future-token leakage.

## 5. Joint training, computation, and empirical evaluation

The overall objective is

$$
\mathcal L_{\mathrm{total}}
=
\mathcal L_{\mathrm{NTP}}
+
\alpha\mathcal L_{\mathrm{NCP}}
+
\beta\mathcal L_{\mathrm{VQ}},
$$

where $\mathcal L_{\mathrm{NTP}}$ is causal token cross-entropy and $\alpha,\beta$ weight the concept-prediction and vector-quantization objectives. The NTP term is

$$
\mathcal L_{\mathrm{NTP}}
=
-\frac{1}{T-1}
\sum_{t=1}^{T-1}
\log
p_{\theta_d}
\left(
x_{t+1}\mid\widetilde{\mathbf h}_{\leq t}
\right).
$$

Matrix parameters use Moonlight Muon, while embeddings, biases, and other non-Muon parameters use AdamW. The default learning rate is $6\times10^{-5}$ with the same cosine scheduler as OLMo-3-7B.

Training used 5.73 trillion tokens from Dolma-3, with Stage 1 using the Dolma 3 Mix and Stage 2 using Dolma 3 Dolmino. The main comparisons were OLMo-3-7B, a 34-block computation-aligned standard Transformer, and a 40-block parameter-aligned standard Transformer. In block notation:

| Model | Parameters | Computation |
|---|---:|---:|
| NCP-ArchPreview | $40P_{\mathrm{blk}}$ | $34F_{\mathrm{blk}}$ |
| Vanilla | $32P_{\mathrm{blk}}$ | $32F_{\mathrm{blk}}$ |
| Size-aligned Vanilla | $40P_{\mathrm{blk}}$ | $40F_{\mathrm{blk}}$ |
| Computation-aligned Vanilla | $34P_{\mathrm{blk}}$ | $34F_{\mathrm{blk}}$ |

The Concept Module contributes roughly eight blocks’ worth of parameters, but operates at one-quarter sequence length. The authors caution that analytical FLOPs omit memory traffic, source-state materialization, reductions, and kernel-launch overhead.

### Pretraining efficiency

Relative to OLMo-3-7B trained on the same data:

- NCP-ArchPreview reached the final OLMo-3-7B loss after 51.3% as many Stage-1 training tokens, corresponding to a reported $1.95\times$ convergence speedup.
- Its Stage-1 loss was 0.091 lower.
- During Stage 2, it reached the final OLMo-3-7B loss after 66.2% of the tokens, corresponding to a reported $1.51\times$ convergence speedup.
- Its final Stage-2 loss was 0.027 lower.
- Scaling-law experiments reported a $1.74\times$ compute-efficiency improvement relative to compute-optimal OLMo-3 training.
- Against the parameter-aligned 40-block baseline, it approached comparable performance while using 85% of the analytical training computation.

### Downstream performance

On a 26-task higher-is-better aggregate, the reported macro averages were:

| Stage | OLMo-3-7B | NCP-ArchPreview | Improvement |
|---|---:|---:|---:|
| Stage 1 | 46.59 | 49.04 | +2.45 |
| Stage 2 | 56.98 | 57.57 | +0.59 |

The strongest Stage-1 gains included MATH average from 20.79 to 24.54, Code average from 25.15 to 27.79, MC-Non-STEM from 70.08 to 74.71, and MMLU average from 54.50 to 56.73.

Stage-1 GSM8K increased from 39.27 to 45.26, a gain of 5.99 points. Other Stage-1 results included GSM-Symbolic from 18.85 to 22.80, Minerva from 12.52 to 15.63, and MATH-500 from 12.52 to 14.48. At Stage 2, GSM8K increased from 79.68 to 83.02 and GSM-Symbolic from 57.32 to 60.32.

The progressive ablations showed the ordering

$$
\text{Vanilla}
\rightarrow
\text{Vanilla+Concept Module}
\rightarrow
\text{Vanilla+Concept Module+Residual}
\rightarrow
\text{Vanilla+Concept Module+Residual+NCP}.
$$

This separates three reported sources of improvement: the compressed latent architecture, hierarchical residual routing, and the explicit NCP objective. The results therefore do not attribute all gains solely to the NCP loss.

## 6. Post-pretraining applications and limitations

The learned concept space is used in two post-pretraining settings: lightweight domain adaptation through VQ updates and concept injection into DFlash2 speculative decoding.

### VQ-only domain adaptation

The reported “+VQ” method freezes the token-level backbone and updates the VQ codebooks and concept-prediction heads. It trains only 17 million parameters without adding parameters to the model. This is compared with full fine-tuning of 8.9 billion parameters and LoRA with 17 million trainable and 17 million added parameters.

For code adaptation, the code average increased from 30.04 to 32.69. For mathematics, the math average increased from 30.56 to 34.83, including GSM8K from 46.78 to 52.69. On TriviaQA, the score increased from 40.28 to 49.47, while the general average changed from 68.68 to 68.71.

With micro-batch size 1 on eight GPUs, reported throughput and memory usage were:

| Method | Throughput | Memory |
|---|---:|---:|
| VQ | 15,632 tokens/s/GPU | 32.4% |
| LoRA | 10,400 tokens/s/GPU | 49.0% |
| Full | 7,739 tokens/s/GPU | 90.3% |

The authors interpret these results as evidence that the concept vocabulary can serve as a lightweight domain-adaptation interface. However, because the token backbone remains frozen, VQ adaptation cannot directly rewrite factual associations stored in that backbone.

### Concept injection into speculative decoding

The concept states were also injected into a DFlash2 block-parallel speculative drafter. For a verified prefix ending at token position $t$, the drafter uses the last completed concept $\mathbf c_{\lfloor t/k\rfloor}$. At drafter layer $\ell$ and proposal position $j$, the modified state is

$$
\widetilde{\mathbf s}_{t,j}^{(\ell)}
=
\mathbf s_{t,j}^{(\ell)}
+
\tanh(g^{(\ell)})
\odot
\operatorname{RMSNorm}_{\ell}
\left(
\mathbf c_{\lfloor t/k\rfloor}
\right).
$$

The modification adds 0.04 million parameters to a 1.1-billion-parameter drafter and requires no additional Target-model computation. The reported mean accepted length improved by 4.17% on average:

| Benchmark | Baseline MAL | Concept MAL | Relative gain |
|---|---:|---:|---:|
| GSM8K | 6.351 | 6.537 | +2.93% |
| MATH | 6.105 | 6.240 | +2.22% |
| HumanEval | 5.432 | 5.845 | +7.59% |
| MBPP | 5.844 | 6.099 | +4.37% |
| Macro average | 5.933 | 6.180 | +4.17% |

### Limitations and unresolved issues

The reported evidence does not establish that product-quantized concepts correspond to uniquely identifiable semantic units. Concepts are fixed-span representations learned from hidden states, and the codebooks are optimized jointly with the language model. Their semantic interpretation may therefore depend on the training distribution, architecture, and quantization configuration.

The architecture also introduces several engineering and methodological qualifications:

- The claimed efficiency advantages depend partly on compressed concept-sequence computation; analytical FLOPs do not include memory traffic, reductions, source-state materialization, or kernel-launch overhead.
- The concept boundaries are fixed at four tokens rather than dynamically inferred.
- Predicted concepts are expected codeword combinations, not necessarily discrete sampled concepts during the differentiable training path.
- The ablations show that performance improvements arise from the latent architecture, hierarchical residuals, and NCP jointly; the NCP objective alone is not isolated as the sole source of the reported gains.
- Stage-2 improvements are smaller and more uneven than Stage-1 improvements, with losses on several code and ARC metrics.
- VQ-only adaptation is efficient but cannot directly modify factual associations encoded in the frozen token backbone.
- The post-pretraining speculative-decoding experiment demonstrates improved mean accepted length, but not a universal improvement across all decoding regimes.
- The technical report presents the largest demonstration of a latent-space language model to date within its reported setting, but this is an architectural and experimental characterization rather than a proof that concept-level prediction is universally superior to token-level training.

NCP-ArchPreview is consequently best characterized as a hierarchical autoregressive language model in which compressed, product-quantized representations are predicted explicitly and returned to the token-generation pathway. Its central architectural claim is that latent concept prediction can provide a useful intermediate predictive structure while preserving ordinary causal decoding. The reported results associate this structure with faster pretraining convergence, improved downstream performance, efficient VQ-based adaptation, and better speculative-drafting acceptance.

Source: https://www.emergentmind.com/topics/ncp-archpreview