NCP-ArchPreview: Latent-Space Language Model
- NCP-ArchPreview is a hierarchical language model that extends causal autoregressive pretraining by jointly predicting both surface tokens and discrete, product-quantized concepts.
- NCP-AchPreview’s architecture consists of a token encoder, a compressed concept module, and a token decoder, enabling competitive model efficiency and faster pretraining convergence.
- Key applications of NCP-ArchPreview include lightweight domain adaptation and speculative decoding in language modeling tasks, with reported gains across multiple benchmarks, including GSM8K and MATH.
NCP-ArchPreview is a latent-space language-model architecture introduced to extend causal autoregressive pretraining beyond standard next-token prediction (NTP). It jointly predicts surface tokens and discrete, product-quantized concepts spanning multiple tokens. The architecture comprises a token encoder, a compressed concept module, and a token decoder; predicted concepts are causally fed back into token-level generation. A technical report describes an 8.9-billion-parameter implementation trained on 5.73 trillion Dolma-3 tokens, together with experiments on pretraining efficiency, downstream performance, lightweight domain adaptation, and speculative decoding (Cao et al., 9 Sep 2026).
1. Motivation and conceptual basis
Standard causal language modeling optimizes the conditional probability of the next token,
through the NTP objective. NCP-ArchPreview retains this objective but adds Next Concept Prediction (NCP), in which the model predicts a discrete latent representation associated with a future span of tokens. The motivation is that hidden states in LLMs already encode abstractions extending across multiple token positions, including entities, propositions, syntactic structures, mathematical substeps, and other semantic regularities. NTP supervises these abstractions only indirectly through individual token likelihoods.
NCP-ArchPreview therefore introduces an explicit predictive hierarchy:
The concept objective is not implemented as a replacement for token prediction. Instead, predicted concepts influence the token decoder through causal feedback, while the final output remains a standard autoregressive token sequence. This distinguishes NCP-ArchPreview from ordinary multi-token prediction: the targets are representations summarizing multiple token positions rather than collections of independent future tokens.
The architecture has three principal computational stages:
- Token Encoder: produces token-resolution hidden states.
- Concept Module: processes a sequence compressed by a factor of four and predicts future quantized concepts.
- Token Decoder: returns to token resolution and uses both token-level states and predicted concepts to produce next-token distributions.
The reported implementation has approximately 8.94 billion parameters and a maximum context length of 8,192 tokens. Its principal dimensions include hidden size 4,096, FFN size 11,008, 32 attention heads, 32 key-value groups, head dimension 128, vocabulary size 100,278, SwiGLU activation, RMSNorm with , RoPE with base 500,000, a token-level local-attention window of 4,096, and zero dropout.
2. Network architecture and hierarchical residual routing
The token-level Transformer consists of 16 encoder layers and 16 decoder layers. Between them is an eight-layer causal Concept Module operating at one-quarter of the token sequence length. The concept module consequently adds approximately eight ordinary Transformer blocks in parameter cost while requiring substantially less full-resolution computation.
The model extends ordinary residual propagation with two mechanisms:
- Intra-Module Residual Connections (IRC) dynamically combine representations from different depths within a module.
- Cross-Module Residual Connections (CRC) transmit representations between the Token Encoder, Concept Module, and Token Decoder.
For module , let denote the input to layer , and let be its block output. Standard residual propagation is
IRC instead forms a set of representations from multiple depths,
and computes mixing coefficients through an MLP:
The next state is
0
The coefficients are not normalized, so signed combinations are possible. Initialization sets the final coefficient to one and all other coefficients to zero, initially recovering standard residual-Transformer behavior.
CRC aligns source representations to the target sequence granularity through repetition or chunk pooling. For source states 1 and target state 2, it computes
3
followed by
4
and
5
where 6 is a learned diagonal scaling vector. The principal CRC paths are Token Encoder 7 Concept Module, Token Encoder 8 Token Decoder, and Concept Module 9 Token Decoder. Cross-module scales are initialized small so that training begins close to the original OLMo-style model.
3. Concept formation and product-quantized vocabulary
For an input sequence 0, the Token Encoder produces hidden states
1
These states are passed toward the decoder, compressed into concepts, and used to learn the concept codebooks.
The compression factor is 2. Every contiguous group of four token states is mean-pooled:
3
where
4
Each concept therefore summarizes a fixed four-token span. The reported model does not use input-adaptive span boundaries; concept granularity is fixed at four tokens.
Each continuous concept is divided into 5 segments:
6
With 7, each segment has dimension 8. Every segment has an independent codebook,
9
The codeword selected for segment 0 of concept 1 is
2
with quantized segment
3
The complete quantized concept is
4
Although each segment codebook has only 128 entries, product quantization provides a latent vocabulary with capacity
5
possible code combinations. The model does not instantiate this as one monolithic table. Instead, it predicts 32 categorical distributions and combines their expected codewords.
The VQ loss is
6
where 7 denotes stop-gradient. This moves selected codebook vectors toward encoder-generated concepts without directly propagating the VQ reconstruction gradient through the Token Encoder. The encoder can nevertheless receive gradients through NCP and NTP.
4. Next Concept Prediction and causal feedback
The Concept Module receives preceding continuous concepts,
8
and produces a latent state
9
For each product-quantization segment 0, a prediction head outputs a distribution over 128 codewords:
1
The predicted segment is the expected codeword,
2
and the complete predicted concept is
3
Expectation over codewords avoids nondifferentiable argmax selection and constrains predictions to combinations of learned codebook entries. The NCP target is the continuous pooled representation, detached from the gradient:
4
An equivalent segment-normalized form divides by 5 and sums segment-wise squared errors. Gradients flow through the Concept Module, prediction heads, codebook-dependent prediction pathway, and preceding concept history, but not through the target concept.
To influence token generation, predicted concepts are repeated four times and shifted forward to preserve causality. For token position 6, the broadcast concept signal is
7
The zero prefix prevents the decoder from using a concept derived from tokens whose prediction it is supervising. The token representation supplied to the decoder is
8
The Token Decoder then models
9
The Concept Module-to-Decoder CRC path uses the same causal shift. Consequently, both direct concept injection and cross-module concept residuals are designed to avoid future-token leakage.
5. Joint training, computation, and empirical evaluation
The overall objective is
0
where 1 is causal token cross-entropy and 2 weight the concept-prediction and vector-quantization objectives. The NTP term is
3
Matrix parameters use Moonlight Muon, while embeddings, biases, and other non-Muon parameters use AdamW. The default learning rate is 4 with the same cosine scheduler as OLMo-3-7B.
Training used 5.73 trillion tokens from Dolma-3, with Stage 1 using the Dolma 3 Mix and Stage 2 using Dolma 3 Dolmino. The main comparisons were OLMo-3-7B, a 34-block computation-aligned standard Transformer, and a 40-block parameter-aligned standard Transformer. In block notation:
| Model | Parameters | Computation |
|---|---|---|
| NCP-ArchPreview | 5 | 6 |
| Vanilla | 7 | 8 |
| Size-aligned Vanilla | 9 | 0 |
| Computation-aligned Vanilla | 1 | 2 |
The Concept Module contributes roughly eight blocks’ worth of parameters, but operates at one-quarter sequence length. The authors caution that analytical FLOPs omit memory traffic, source-state materialization, reductions, and kernel-launch overhead.
Pretraining efficiency
Relative to OLMo-3-7B trained on the same data:
- NCP-ArchPreview reached the final OLMo-3-7B loss after 51.3% as many Stage-1 training tokens, corresponding to a reported 3 convergence speedup.
- Its Stage-1 loss was 0.091 lower.
- During Stage 2, it reached the final OLMo-3-7B loss after 66.2% of the tokens, corresponding to a reported 4 convergence speedup.
- Its final Stage-2 loss was 0.027 lower.
- Scaling-law experiments reported a 5 compute-efficiency improvement relative to compute-optimal OLMo-3 training.
- Against the parameter-aligned 40-block baseline, it approached comparable performance while using 85% of the analytical training computation.
Downstream performance
On a 26-task higher-is-better aggregate, the reported macro averages were:
| Stage | OLMo-3-7B | NCP-ArchPreview | Improvement |
|---|---|---|---|
| Stage 1 | 46.59 | 49.04 | +2.45 |
| Stage 2 | 56.98 | 57.57 | +0.59 |
The strongest Stage-1 gains included MATH average from 20.79 to 24.54, Code average from 25.15 to 27.79, MC-Non-STEM from 70.08 to 74.71, and MMLU average from 54.50 to 56.73.
Stage-1 GSM8K increased from 39.27 to 45.26, a gain of 5.99 points. Other Stage-1 results included GSM-Symbolic from 18.85 to 22.80, Minerva from 12.52 to 15.63, and MATH-500 from 12.52 to 14.48. At Stage 2, GSM8K increased from 79.68 to 83.02 and GSM-Symbolic from 57.32 to 60.32.
The progressive ablations showed the ordering
6
This separates three reported sources of improvement: the compressed latent architecture, hierarchical residual routing, and the explicit NCP objective. The results therefore do not attribute all gains solely to the NCP loss.
6. Post-pretraining applications and limitations
The learned concept space is used in two post-pretraining settings: lightweight domain adaptation through VQ updates and concept injection into DFlash2 speculative decoding.
VQ-only domain adaptation
The reported “+VQ” method freezes the token-level backbone and updates the VQ codebooks and concept-prediction heads. It trains only 17 million parameters without adding parameters to the model. This is compared with full fine-tuning of 8.9 billion parameters and LoRA with 17 million trainable and 17 million added parameters.
For code adaptation, the code average increased from 30.04 to 32.69. For mathematics, the math average increased from 30.56 to 34.83, including GSM8K from 46.78 to 52.69. On TriviaQA, the score increased from 40.28 to 49.47, while the general average changed from 68.68 to 68.71.
With micro-batch size 1 on eight GPUs, reported throughput and memory usage were:
| Method | Throughput | Memory |
|---|---|---|
| VQ | 15,632 tokens/s/GPU | 32.4% |
| LoRA | 10,400 tokens/s/GPU | 49.0% |
| Full | 7,739 tokens/s/GPU | 90.3% |
The authors interpret these results as evidence that the concept vocabulary can serve as a lightweight domain-adaptation interface. However, because the token backbone remains frozen, VQ adaptation cannot directly rewrite factual associations stored in that backbone.
Concept injection into speculative decoding
The concept states were also injected into a DFlash2 block-parallel speculative drafter. For a verified prefix ending at token position 7, the drafter uses the last completed concept 8. At drafter layer 9 and proposal position 0, the modified state is
1
The modification adds 0.04 million parameters to a 1.1-billion-parameter drafter and requires no additional Target-model computation. The reported mean accepted length improved by 4.17% on average:
| Benchmark | Baseline MAL | Concept MAL | Relative gain |
|---|---|---|---|
| GSM8K | 6.351 | 6.537 | +2.93% |
| MATH | 6.105 | 6.240 | +2.22% |
| HumanEval | 5.432 | 5.845 | +7.59% |
| MBPP | 5.844 | 6.099 | +4.37% |
| Macro average | 5.933 | 6.180 | +4.17% |
Limitations and unresolved issues
The reported evidence does not establish that product-quantized concepts correspond to uniquely identifiable semantic units. Concepts are fixed-span representations learned from hidden states, and the codebooks are optimized jointly with the LLM. Their semantic interpretation may therefore depend on the training distribution, architecture, and quantization configuration.
The architecture also introduces several engineering and methodological qualifications:
- The claimed efficiency advantages depend partly on compressed concept-sequence computation; analytical FLOPs do not include memory traffic, reductions, source-state materialization, or kernel-launch overhead.
- The concept boundaries are fixed at four tokens rather than dynamically inferred.
- Predicted concepts are expected codeword combinations, not necessarily discrete sampled concepts during the differentiable training path.
- The ablations show that performance improvements arise from the latent architecture, hierarchical residuals, and NCP jointly; the NCP objective alone is not isolated as the sole source of the reported gains.
- Stage-2 improvements are smaller and more uneven than Stage-1 improvements, with losses on several code and ARC metrics.
- VQ-only adaptation is efficient but cannot directly modify factual associations encoded in the frozen token backbone.
- The post-pretraining speculative-decoding experiment demonstrates improved mean accepted length, but not a universal improvement across all decoding regimes.
- The technical report presents the largest demonstration of a latent-space LLM to date within its reported setting, but this is an architectural and experimental characterization rather than a proof that concept-level prediction is universally superior to token-level training.
NCP-ArchPreview is consequently best characterized as a hierarchical autoregressive LLM in which compressed, product-quantized representations are predicted explicitly and returned to the token-generation pathway. Its central architectural claim is that latent concept prediction can provide a useful intermediate predictive structure while preserving ordinary causal decoding. The reported results associate this structure with faster pretraining convergence, improved downstream performance, efficient VQ-based adaptation, and better speculative-drafting acceptance.