Papers
Topics
Authors
Recent
Search
2000 character limit reached

NCP-ArchPreview: Latent-Space Language Model

Updated 14 September 2026
  • NCP-ArchPreview is a hierarchical language model that extends causal autoregressive pretraining by jointly predicting both surface tokens and discrete, product-quantized concepts.
  • NCP-AchPreview’s architecture consists of a token encoder, a compressed concept module, and a token decoder, enabling competitive model efficiency and faster pretraining convergence.
  • Key applications of NCP-ArchPreview include lightweight domain adaptation and speculative decoding in language modeling tasks, with reported gains across multiple benchmarks, including GSM8K and MATH.

NCP-ArchPreview is a latent-space language-model architecture introduced to extend causal autoregressive pretraining beyond standard next-token prediction (NTP). It jointly predicts surface tokens and discrete, product-quantized concepts spanning multiple tokens. The architecture comprises a token encoder, a compressed concept module, and a token decoder; predicted concepts are causally fed back into token-level generation. A technical report describes an 8.9-billion-parameter implementation trained on 5.73 trillion Dolma-3 tokens, together with experiments on pretraining efficiency, downstream performance, lightweight domain adaptation, and speculative decoding (Cao et al., 9 Sep 2026).

1. Motivation and conceptual basis

Standard causal language modeling optimizes the conditional probability of the next token,

p(xt+1∣x≤t),p(x_{t+1}\mid x_{\leq t}),

through the NTP objective. NCP-ArchPreview retains this objective but adds Next Concept Prediction (NCP), in which the model predicts a discrete latent representation associated with a future span of tokens. The motivation is that hidden states in LLMs already encode abstractions extending across multiple token positions, including entities, propositions, syntactic structures, mathematical substeps, and other semantic regularities. NTP supervises these abstractions only indirectly through individual token likelihoods.

NCP-ArchPreview therefore introduces an explicit predictive hierarchy:

tokens→continuous concepts→discrete concept vocabulary→predicted future concepts→token generation.\text{tokens} \rightarrow \text{continuous concepts} \rightarrow \text{discrete concept vocabulary} \rightarrow \text{predicted future concepts} \rightarrow \text{token generation}.

The concept objective is not implemented as a replacement for token prediction. Instead, predicted concepts influence the token decoder through causal feedback, while the final output remains a standard autoregressive token sequence. This distinguishes NCP-ArchPreview from ordinary multi-token prediction: the targets are representations summarizing multiple token positions rather than collections of independent future tokens.

The architecture has three principal computational stages:

  1. Token Encoder: produces token-resolution hidden states.
  2. Concept Module: processes a sequence compressed by a factor of four and predicts future quantized concepts.
  3. Token Decoder: returns to token resolution and uses both token-level states and predicted concepts to produce next-token distributions.

The reported implementation has approximately 8.94 billion parameters and a maximum context length of 8,192 tokens. Its principal dimensions include hidden size 4,096, FFN size 11,008, 32 attention heads, 32 key-value groups, head dimension 128, vocabulary size 100,278, SwiGLU activation, RMSNorm with ϵ=10−6\epsilon=10^{-6}, RoPE with base 500,000, a token-level local-attention window of 4,096, and zero dropout.

2. Network architecture and hierarchical residual routing

The token-level Transformer consists of 16 encoder layers and 16 decoder layers. Between them is an eight-layer causal Concept Module operating at one-quarter of the token sequence length. The concept module consequently adds approximately eight ordinary Transformer blocks in parameter cost while requiring substantially less full-resolution computation.

The model extends ordinary residual propagation with two mechanisms:

  • Intra-Module Residual Connections (IRC) dynamically combine representations from different depths within a module.
  • Cross-Module Residual Connections (CRC) transmit representations between the Token Encoder, Concept Module, and Token Decoder.

For module ss, let Hℓs\mathbf H_\ell^s denote the input to layer ℓ\ell, and let Rℓs=Fℓs(Hℓs)\mathbf R_\ell^s=F_\ell^s(\mathbf H_\ell^s) be its block output. Standard residual propagation is

Hℓ+1s=Hℓs+Rℓs.\mathbf H_{\ell+1}^s=\mathbf H_\ell^s+\mathbf R_\ell^s.

IRC instead forms a set of representations from multiple depths,

Xℓs={H1s, H1s+R1s, …, Hℓs+Rℓs},\mathcal X_\ell^s= \left\{ \mathbf H_1^s,\, \mathbf H_1^s+\mathbf R_1^s,\, \ldots,\, \mathbf H_\ell^s+\mathbf R_\ell^s \right\},

and computes mixing coefficients through an MLP:

wℓs=MLP⁡ℓs(Rℓs).\mathbf w_\ell^s=\operatorname{MLP}_\ell^s(\mathbf R_\ell^s).

The next state is

tokens→continuous concepts→discrete concept vocabulary→predicted future concepts→token generation.\text{tokens} \rightarrow \text{continuous concepts} \rightarrow \text{discrete concept vocabulary} \rightarrow \text{predicted future concepts} \rightarrow \text{token generation}.0

The coefficients are not normalized, so signed combinations are possible. Initialization sets the final coefficient to one and all other coefficients to zero, initially recovering standard residual-Transformer behavior.

CRC aligns source representations to the target sequence granularity through repetition or chunk pooling. For source states tokens→continuous concepts→discrete concept vocabulary→predicted future concepts→token generation.\text{tokens} \rightarrow \text{continuous concepts} \rightarrow \text{discrete concept vocabulary} \rightarrow \text{predicted future concepts} \rightarrow \text{token generation}.1 and target state tokens→continuous concepts→discrete concept vocabulary→predicted future concepts→token generation.\text{tokens} \rightarrow \text{continuous concepts} \rightarrow \text{discrete concept vocabulary} \rightarrow \text{predicted future concepts} \rightarrow \text{token generation}.2, it computes

tokens→continuous concepts→discrete concept vocabulary→predicted future concepts→token generation.\text{tokens} \rightarrow \text{continuous concepts} \rightarrow \text{discrete concept vocabulary} \rightarrow \text{predicted future concepts} \rightarrow \text{token generation}.3

followed by

tokens→continuous concepts→discrete concept vocabulary→predicted future concepts→token generation.\text{tokens} \rightarrow \text{continuous concepts} \rightarrow \text{discrete concept vocabulary} \rightarrow \text{predicted future concepts} \rightarrow \text{token generation}.4

and

tokens→continuous concepts→discrete concept vocabulary→predicted future concepts→token generation.\text{tokens} \rightarrow \text{continuous concepts} \rightarrow \text{discrete concept vocabulary} \rightarrow \text{predicted future concepts} \rightarrow \text{token generation}.5

where tokens→continuous concepts→discrete concept vocabulary→predicted future concepts→token generation.\text{tokens} \rightarrow \text{continuous concepts} \rightarrow \text{discrete concept vocabulary} \rightarrow \text{predicted future concepts} \rightarrow \text{token generation}.6 is a learned diagonal scaling vector. The principal CRC paths are Token Encoder tokens→continuous concepts→discrete concept vocabulary→predicted future concepts→token generation.\text{tokens} \rightarrow \text{continuous concepts} \rightarrow \text{discrete concept vocabulary} \rightarrow \text{predicted future concepts} \rightarrow \text{token generation}.7 Concept Module, Token Encoder tokens→continuous concepts→discrete concept vocabulary→predicted future concepts→token generation.\text{tokens} \rightarrow \text{continuous concepts} \rightarrow \text{discrete concept vocabulary} \rightarrow \text{predicted future concepts} \rightarrow \text{token generation}.8 Token Decoder, and Concept Module tokens→continuous concepts→discrete concept vocabulary→predicted future concepts→token generation.\text{tokens} \rightarrow \text{continuous concepts} \rightarrow \text{discrete concept vocabulary} \rightarrow \text{predicted future concepts} \rightarrow \text{token generation}.9 Token Decoder. Cross-module scales are initialized small so that training begins close to the original OLMo-style model.

3. Concept formation and product-quantized vocabulary

For an input sequence ϵ=10−6\epsilon=10^{-6}0, the Token Encoder produces hidden states

ϵ=10−6\epsilon=10^{-6}1

These states are passed toward the decoder, compressed into concepts, and used to learn the concept codebooks.

The compression factor is ϵ=10−6\epsilon=10^{-6}2. Every contiguous group of four token states is mean-pooled:

ϵ=10−6\epsilon=10^{-6}3

where

ϵ=10−6\epsilon=10^{-6}4

Each concept therefore summarizes a fixed four-token span. The reported model does not use input-adaptive span boundaries; concept granularity is fixed at four tokens.

Each continuous concept is divided into ϵ=10−6\epsilon=10^{-6}5 segments:

ϵ=10−6\epsilon=10^{-6}6

With ϵ=10−6\epsilon=10^{-6}7, each segment has dimension ϵ=10−6\epsilon=10^{-6}8. Every segment has an independent codebook,

ϵ=10−6\epsilon=10^{-6}9

The codeword selected for segment ss0 of concept ss1 is

ss2

with quantized segment

ss3

The complete quantized concept is

ss4

Although each segment codebook has only 128 entries, product quantization provides a latent vocabulary with capacity

ss5

possible code combinations. The model does not instantiate this as one monolithic table. Instead, it predicts 32 categorical distributions and combines their expected codewords.

The VQ loss is

ss6

where ss7 denotes stop-gradient. This moves selected codebook vectors toward encoder-generated concepts without directly propagating the VQ reconstruction gradient through the Token Encoder. The encoder can nevertheless receive gradients through NCP and NTP.

4. Next Concept Prediction and causal feedback

The Concept Module receives preceding continuous concepts,

ss8

and produces a latent state

ss9

For each product-quantization segment Hℓs\mathbf H_\ell^s0, a prediction head outputs a distribution over 128 codewords:

Hℓs\mathbf H_\ell^s1

The predicted segment is the expected codeword,

Hℓs\mathbf H_\ell^s2

and the complete predicted concept is

Hℓs\mathbf H_\ell^s3

Expectation over codewords avoids nondifferentiable argmax selection and constrains predictions to combinations of learned codebook entries. The NCP target is the continuous pooled representation, detached from the gradient:

Hℓs\mathbf H_\ell^s4

An equivalent segment-normalized form divides by Hℓs\mathbf H_\ell^s5 and sums segment-wise squared errors. Gradients flow through the Concept Module, prediction heads, codebook-dependent prediction pathway, and preceding concept history, but not through the target concept.

To influence token generation, predicted concepts are repeated four times and shifted forward to preserve causality. For token position Hℓs\mathbf H_\ell^s6, the broadcast concept signal is

Hℓs\mathbf H_\ell^s7

The zero prefix prevents the decoder from using a concept derived from tokens whose prediction it is supervising. The token representation supplied to the decoder is

Hℓs\mathbf H_\ell^s8

The Token Decoder then models

Hℓs\mathbf H_\ell^s9

The Concept Module-to-Decoder CRC path uses the same causal shift. Consequently, both direct concept injection and cross-module concept residuals are designed to avoid future-token leakage.

5. Joint training, computation, and empirical evaluation

The overall objective is

ℓ\ell0

where ℓ\ell1 is causal token cross-entropy and ℓ\ell2 weight the concept-prediction and vector-quantization objectives. The NTP term is

ℓ\ell3

Matrix parameters use Moonlight Muon, while embeddings, biases, and other non-Muon parameters use AdamW. The default learning rate is ℓ\ell4 with the same cosine scheduler as OLMo-3-7B.

Training used 5.73 trillion tokens from Dolma-3, with Stage 1 using the Dolma 3 Mix and Stage 2 using Dolma 3 Dolmino. The main comparisons were OLMo-3-7B, a 34-block computation-aligned standard Transformer, and a 40-block parameter-aligned standard Transformer. In block notation:

Model Parameters Computation
NCP-ArchPreview ℓ\ell5 ℓ\ell6
Vanilla ℓ\ell7 ℓ\ell8
Size-aligned Vanilla ℓ\ell9 Rℓs=Fℓs(Hℓs)\mathbf R_\ell^s=F_\ell^s(\mathbf H_\ell^s)0
Computation-aligned Vanilla Rℓs=Fℓs(Hℓs)\mathbf R_\ell^s=F_\ell^s(\mathbf H_\ell^s)1 Rℓs=Fℓs(Hℓs)\mathbf R_\ell^s=F_\ell^s(\mathbf H_\ell^s)2

The Concept Module contributes roughly eight blocks’ worth of parameters, but operates at one-quarter sequence length. The authors caution that analytical FLOPs omit memory traffic, source-state materialization, reductions, and kernel-launch overhead.

Pretraining efficiency

Relative to OLMo-3-7B trained on the same data:

  • NCP-ArchPreview reached the final OLMo-3-7B loss after 51.3% as many Stage-1 training tokens, corresponding to a reported Rℓs=Fℓs(Hℓs)\mathbf R_\ell^s=F_\ell^s(\mathbf H_\ell^s)3 convergence speedup.
  • Its Stage-1 loss was 0.091 lower.
  • During Stage 2, it reached the final OLMo-3-7B loss after 66.2% of the tokens, corresponding to a reported Rℓs=Fℓs(Hℓs)\mathbf R_\ell^s=F_\ell^s(\mathbf H_\ell^s)4 convergence speedup.
  • Its final Stage-2 loss was 0.027 lower.
  • Scaling-law experiments reported a Rℓs=Fℓs(Hℓs)\mathbf R_\ell^s=F_\ell^s(\mathbf H_\ell^s)5 compute-efficiency improvement relative to compute-optimal OLMo-3 training.
  • Against the parameter-aligned 40-block baseline, it approached comparable performance while using 85% of the analytical training computation.

Downstream performance

On a 26-task higher-is-better aggregate, the reported macro averages were:

Stage OLMo-3-7B NCP-ArchPreview Improvement
Stage 1 46.59 49.04 +2.45
Stage 2 56.98 57.57 +0.59

The strongest Stage-1 gains included MATH average from 20.79 to 24.54, Code average from 25.15 to 27.79, MC-Non-STEM from 70.08 to 74.71, and MMLU average from 54.50 to 56.73.

Stage-1 GSM8K increased from 39.27 to 45.26, a gain of 5.99 points. Other Stage-1 results included GSM-Symbolic from 18.85 to 22.80, Minerva from 12.52 to 15.63, and MATH-500 from 12.52 to 14.48. At Stage 2, GSM8K increased from 79.68 to 83.02 and GSM-Symbolic from 57.32 to 60.32.

The progressive ablations showed the ordering

Rℓs=Fℓs(Hℓs)\mathbf R_\ell^s=F_\ell^s(\mathbf H_\ell^s)6

This separates three reported sources of improvement: the compressed latent architecture, hierarchical residual routing, and the explicit NCP objective. The results therefore do not attribute all gains solely to the NCP loss.

6. Post-pretraining applications and limitations

The learned concept space is used in two post-pretraining settings: lightweight domain adaptation through VQ updates and concept injection into DFlash2 speculative decoding.

VQ-only domain adaptation

The reported “+VQ” method freezes the token-level backbone and updates the VQ codebooks and concept-prediction heads. It trains only 17 million parameters without adding parameters to the model. This is compared with full fine-tuning of 8.9 billion parameters and LoRA with 17 million trainable and 17 million added parameters.

For code adaptation, the code average increased from 30.04 to 32.69. For mathematics, the math average increased from 30.56 to 34.83, including GSM8K from 46.78 to 52.69. On TriviaQA, the score increased from 40.28 to 49.47, while the general average changed from 68.68 to 68.71.

With micro-batch size 1 on eight GPUs, reported throughput and memory usage were:

Method Throughput Memory
VQ 15,632 tokens/s/GPU 32.4%
LoRA 10,400 tokens/s/GPU 49.0%
Full 7,739 tokens/s/GPU 90.3%

The authors interpret these results as evidence that the concept vocabulary can serve as a lightweight domain-adaptation interface. However, because the token backbone remains frozen, VQ adaptation cannot directly rewrite factual associations stored in that backbone.

Concept injection into speculative decoding

The concept states were also injected into a DFlash2 block-parallel speculative drafter. For a verified prefix ending at token position Rℓs=Fℓs(Hℓs)\mathbf R_\ell^s=F_\ell^s(\mathbf H_\ell^s)7, the drafter uses the last completed concept Rℓs=Fℓs(Hℓs)\mathbf R_\ell^s=F_\ell^s(\mathbf H_\ell^s)8. At drafter layer Rℓs=Fℓs(Hℓs)\mathbf R_\ell^s=F_\ell^s(\mathbf H_\ell^s)9 and proposal position Hℓ+1s=Hℓs+Rℓs.\mathbf H_{\ell+1}^s=\mathbf H_\ell^s+\mathbf R_\ell^s.0, the modified state is

Hℓ+1s=Hℓs+Rℓs.\mathbf H_{\ell+1}^s=\mathbf H_\ell^s+\mathbf R_\ell^s.1

The modification adds 0.04 million parameters to a 1.1-billion-parameter drafter and requires no additional Target-model computation. The reported mean accepted length improved by 4.17% on average:

Benchmark Baseline MAL Concept MAL Relative gain
GSM8K 6.351 6.537 +2.93%
MATH 6.105 6.240 +2.22%
HumanEval 5.432 5.845 +7.59%
MBPP 5.844 6.099 +4.37%
Macro average 5.933 6.180 +4.17%

Limitations and unresolved issues

The reported evidence does not establish that product-quantized concepts correspond to uniquely identifiable semantic units. Concepts are fixed-span representations learned from hidden states, and the codebooks are optimized jointly with the LLM. Their semantic interpretation may therefore depend on the training distribution, architecture, and quantization configuration.

The architecture also introduces several engineering and methodological qualifications:

  • The claimed efficiency advantages depend partly on compressed concept-sequence computation; analytical FLOPs do not include memory traffic, reductions, source-state materialization, or kernel-launch overhead.
  • The concept boundaries are fixed at four tokens rather than dynamically inferred.
  • Predicted concepts are expected codeword combinations, not necessarily discrete sampled concepts during the differentiable training path.
  • The ablations show that performance improvements arise from the latent architecture, hierarchical residuals, and NCP jointly; the NCP objective alone is not isolated as the sole source of the reported gains.
  • Stage-2 improvements are smaller and more uneven than Stage-1 improvements, with losses on several code and ARC metrics.
  • VQ-only adaptation is efficient but cannot directly modify factual associations encoded in the frozen token backbone.
  • The post-pretraining speculative-decoding experiment demonstrates improved mean accepted length, but not a universal improvement across all decoding regimes.
  • The technical report presents the largest demonstration of a latent-space LLM to date within its reported setting, but this is an architectural and experimental characterization rather than a proof that concept-level prediction is universally superior to token-level training.

NCP-ArchPreview is consequently best characterized as a hierarchical autoregressive LLM in which compressed, product-quantized representations are predicted explicitly and returned to the token-generation pathway. Its central architectural claim is that latent concept prediction can provide a useful intermediate predictive structure while preserving ordinary causal decoding. The reported results associate this structure with faster pretraining convergence, improved downstream performance, efficient VQ-based adaptation, and better speculative-drafting acceptance.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to NCP-ArchPreview.