Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior
Abstract: Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during LLM pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at larger scales and under more realistic PT data mixtures that combine diverse sources (e.g., code and math). We therefore present a comprehensive study on PPT spanning five PPT tasks, four PT data mixtures, four parameter scales (500M to 7B), and PT budgets of up to 100B tokens. Our results demonstrate that the downstream performance and token efficiency gains of PPT persist at scale, e.g., saving at least 21B PT tokens at the 3B scale. However, in contrast to prior work, we find no consistent evidence that these gains stem from a grammatical prior. Downstream performance does not consistently align with grammatical acceptability across model sizes. Instead, we find that downstream gains arise from PPT tasks that improve long-range retrieval. Finally, PPT performance gains are robust to how PT data mixtures are composed and diminish only when web text is absent. Overall, PPT is a low-cost addition to PT, and future PPT task design should target long-range retrieval rather than natural language grammar.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper studies a technique called synthetic pre-pretraining, or PPT.
Usually, a LLM learns by reading huge amounts of text, such as websites, books, code, and mathematics. In PPT, the model first practices on specially created data that is not normal language. For example, it might practice matching different kinds of brackets:
1 |
( [ ( ] ) ) |
After this short practice phase, the model is trained normally on real text.
Earlier research suggested that this practice helps because the model learns a kind of grammar instinct. The authors of this paper ask whether that explanation is correct—and whether PPT still helps larger models trained on much more data.
2. What questions did the researchers ask?
The paper focuses on several main questions:
- Does PPT still help when LLMs become much larger?
- Does the benefit remain after the model studies tens of billions of tokens?
- Does PPT help when the training data includes not only web text, but also code and mathematics?
- Does PPT improve the model’s understanding of grammar?
- Or does it help for another reason, such as remembering information from far earlier in a passage?
- Which kinds of synthetic practice tasks are useful?
The researchers also wanted to know whether PPT is worth using in real language-model training.
3. How did the researchers investigate this?
Different model sizes
They trained models with approximately:
- 500 million parameters
- 1 billion parameters
- 3 billion parameters
- 7 billion parameters
A parameter is a number inside a model that helps it make predictions. More parameters are like giving a student a larger brain or more notebook space, although a larger model also needs more training.
Different training data
The models were trained on four kinds of data mixtures:
- Mostly web text
- Web text mixed with code
- Web text mixed with mathematics
- A mixture containing academic and science-related material
This allowed the researchers to test whether code and mathematics made PPT unnecessary.
Different synthetic practice tasks
The researchers tested five types of PPT data. Some involved formal grammar, while others involved different kinds of structure.
For example:
- Bracket matching: The model must connect each closing bracket to the correct opening bracket.
- Marked structures: Extra symbols show which brackets belong together.
- Neural cellular automata: The model predicts how a grid changes over time according to fixed rules.
- Set task: The model checks whether a symbol has appeared before.
- Text control: The model gets a short warm-up using real text instead of synthetic data.
The text control was important. It showed whether any improvement came simply from doing 500 extra training steps, rather than from the special synthetic data.
Testing the models
The researchers tested the models in several ways:
- General tasks: reading comprehension, science questions, commonsense reasoning, and language prediction.
- Grammar tests: The BLiMP benchmark compared sentences to see whether the model preferred grammatically correct ones.
- Long-range retrieval tests: The model had to remember and repeat information that appeared much earlier in a long passage.
Long-range retrieval is like reading a whole story and then remembering the name mentioned several pages earlier.
4. What did they find?
PPT continued to help larger models
The main result was that PPT remained useful even for larger models and longer training.
At the 3-billion-parameter size, a model using PPT could reach a certain performance level after about 63 billion training tokens, while a model without PPT needed about 84 billion tokens to reach a similar level.
That means PPT saved at least 21 billion training tokens in that experiment. Since training large models is extremely expensive, this could be important.
The benefit also remained positive for the 7-billion-parameter model, although it became somewhat smaller.
The benefit was not consistently caused by better grammar
The researchers did not find strong, consistent evidence for the earlier explanation that PPT teaches a general grammar skill.
PPT sometimes improved grammar-test scores, but these improvements were inconsistent. In some cases, grammar performance even became worse. The grammar results also did not match the general improvements very well.
For example, a model could perform better at reading comprehension and science questions without showing a reliable improvement on grammar tests.
This suggests that the model was gaining something other than a general “grammar instinct.”
PPT improved long-range retrieval
The strongest explanation was that useful PPT tasks teach the model to find and use information from earlier in a sequence.
For example, in the bracket task, the model has to remember an opening bracket and connect it with the correct closing bracket, even when many other symbols appear between them.
This is similar to tasks such as:
- Finding information earlier in a reading passage
- Remembering a person’s name from the beginning of a story
- Using an earlier fact to answer a question later
- Copying a list that appeared many words ago
The models that improved most on long-range retrieval also tended to improve on downstream tasks such as reading comprehension and language modeling.
Not every synthetic task worked
Three tasks worked fairly well:
- Bracket matching
- Marked structural sequences
- Neural cellular automata
However, the Set task performed poorly. It mainly required the model to remember whether a symbol had appeared before, not to find the exact earlier position where it appeared.
This difference was important. The results suggest that PPT works best when the model must retrieve information from a specific earlier location, rather than merely keep a general list of what it has seen.
Code and mathematics did not replace PPT
The researchers expected that code and mathematics might already teach models about structure, making PPT less useful.
That did not happen. Even when the amount of mathematics was increased thirteenfold, PPT continued to provide similar benefits.
However, the benefit became much smaller when all web text was removed. This suggests that PPT is especially helpful before training on natural web language, even when the training mixture also includes code and mathematics.
5. Why are these findings important?
Training LLMs requires enormous amounts of computing power, time, and electricity. If a short and cheap synthetic warm-up can help a model learn more efficiently, it could reduce the amount of expensive training needed.
The paper also changes how researchers should think about PPT. The important lesson may not be:
“Teach the model grammar before teaching it language.”
Instead, it may be:
“Teach the model to retrieve information from far earlier in a sequence.”
This gives researchers a clearer way to design future synthetic training tasks.
Conclusion
The paper shows that synthetic pre-pretraining is still useful at large scale. It can improve language-model performance and reduce the amount of ordinary training needed, even for models with billions of parameters and training budgets as large as 100 billion tokens.
However, the improvements do not seem to come mainly from learning a general grammar ability. They appear to come from practicing long-range retrieval—remembering and finding important information from earlier in a sequence.
In the future, researchers could create synthetic tasks that directly practice this skill. This might lead to LLMs that learn more efficiently and perform better on tasks involving long passages, reading comprehension, reasoning, and memory.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Causal mechanism of PPT remains unresolved: The paper associates successful PPT with long-range retrieval, but does not establish whether retrieval is the causal mechanism or merely correlated with another property, such as optimization geometry, attention formation, positional encoding use, or reduced loss landscape difficulty.
- The retrieval hypothesis is tested with limited diagnostics: Verbatim retrieval is the primary direct measure, while several downstream benchmarks also require contextual information. More controlled evaluations are needed to distinguish retrieval of an exact earlier span from reasoning, entity tracking, copying, induction, and general context utilization.
- The boundary between effective and harmful PPT tasks is unclear:
k-Shuffle Dyck, MP-Struct Core, and NCA improve performance, whereas Set causes substantial degradation, but the paper does not isolate which task properties explain this contrast—positional specificity, sequence entropy, repetition, difficulty, token distribution, or optimization dynamics. - No systematic task-property ablation is provided: Future work should vary retrieval distance, ambiguity, nesting depth, repetition rate, sequence length, and structural complexity independently to determine which synthetic-task characteristics produce transfer.
- The grammatical-prior hypothesis is not fully ruled out: BLiMP shows inconsistent gains, but the study does not test other forms of syntactic or hierarchical competence, such as longer-range agreement, syntactic generalization to unseen constructions, probing of learned representations, or controlled grammar-learning tasks.
- The apparent conflict with prior grammatical-prior results is unexplained: Differences in tokenizer, architecture, optimizer, batch size, learning-rate schedule, PPT duration, data generation, evaluation protocol, or random seeds are not isolated, so it remains unclear why this study fails to reproduce earlier BLiMP improvements.
- The role of web text is underdetermined: Removing web text reduces the PPT benefit, but the experiments do not identify whether web text is necessary because of its linguistic structure, diversity, quality, scale, deduplication properties, or interaction with the tokenizer.
- The OLMo3 anomaly remains unexplained: OLMo3 shows weak or absent PPT gains at multiple scales, yet the paper does not directly ablate its OCR-derived academic text, filtering procedures, source diversity, document lengths, or domain proportions to identify the responsible factor.
- PT mixture effects are only partially explored: The study compares a small number of predefined mixtures and one manually modified Marin mixture; it does not provide a factorial analysis over natural-language, code, mathematics, academic text, quality, and filtering proportions.
- Code and mathematics may be confounded with data quality and distribution: The conclusion that code and mathematics do not make PPT redundant is based mainly on mixture-level comparisons, where domain content covaries with filtering, source, document format, and quality.
- The 7B scaling result is narrow: At 7B parameters, only
k-Shuffle Dyck and the Marin mixture are evaluated, with a shorter effective training budget than at 3B. It is therefore unclear whether PPT remains useful across other mixtures, PPT tasks, model architectures, and compute-optimal 7B training regimes. - Scaling beyond 7B is untested: The paper cannot determine whether PPT gains eventually disappear, increase, or change mechanism at larger model sizes representative of current frontier systems.
- Long-horizon conclusions are limited to 100B tokens and one main task: Extended training is evaluated primarily with
k-Shuffle Dyck at 3B, with only selected 7B runs. The persistence of gains over longer horizons, additional curricula, or multi-stage training remains unknown. - Token-efficiency claims are not fully compute-normalized: The reported savings in PT tokens do not comprehensively account for PPT generation, preprocessing, training, memory, hardware utilization, optimizer-state resets, and evaluation costs.
- PPT budget is not systematically varied: All main PPT runs use 500 steps, leaving unresolved whether the observed gains depend on this specific duration or whether there is an optimal PPT budget relative to model size, PT budget, and synthetic-task difficulty.
- Optimizer and scheduler reset effects are not isolated: PPT checkpoints are transferred while optimizer moments and the scheduler are reset. The contribution of the transferred weights versus optimizer-state initialization, learning-rate discontinuities, or alternative continuation strategies remains unknown.
- Learning-rate differences introduce a scale confound: The 7B model uses a lower learning rate, so differences between 3B and 7B cannot be attributed solely to parameter count.
- Random-seed robustness is limited: Most configurations are represented by a single run, with three seeds only for a restricted subset of 3B Marin experiments. The statistical reliability of small gains, losses, and mixture-specific anomalies is therefore uncertain.
- The stability criterion is ad hoc: Calling a gain stable when its mean exceeds its checkpoint standard deviation does not constitute a formal significance or uncertainty analysis and may not account for correlated checkpoints or multiple comparisons.
- Checkpoint averaging may obscure training dynamics: Reporting means over steps 5K–10K can conceal when PPT benefits emerge, peak, disappear, or reverse, limiting understanding of whether the gains are initialization effects or persistent changes in learning dynamics.
- Architecture generality is untested: All models use the SmolLM3 architecture and tokenizer, so the findings may not transfer to Llama-style, mixture-of-experts, recurrent, state-space, multimodal, or alternative positional-encoding architectures.
- Tokenizer dependence is unresolved: Synthetic PPT tasks are trained with the natural-language tokenizer, but the effects of vocabulary size, token fragmentation, special-token handling, and tokenizer reinitialization are not examined.
- The evaluation suite may not represent broad capabilities: The downstream benchmarks are mostly small, conventional, zero-shot or few-shot English tasks. Effects on generation quality, instruction following, coding, mathematical reasoning, multilinguality, factual knowledge, calibration, safety, and long-context reasoning remain unknown.
- Benchmark-level gains may reflect task overlap: Some evaluated tasks emphasize contextual retrieval, which aligns with the proposed mechanism. It remains unclear whether PPT improves capabilities unrelated to retrieval or whether the reported average is disproportionately driven by retrieval-heavy benchmarks.
- Long-context generalization is not directly evaluated: The verbatim retrieval setup and the 4,096-token training context do not establish whether PPT improves retrieval at substantially longer contexts, under distractors, across multiple relevant spans, or with degraded positional cues.
- Generalization beyond the synthetic distributions is uncertain: The study does not test whether retrieval improvements transfer across sequence lengths, symbol alphabets, structural rules, noise levels, or synthetic tasks not seen during PPT.
- Potential data and benchmark contamination is not discussed in depth: The relationship between PT mixtures, PPT controls, public benchmark data, and synthetic generation procedures is not sufficiently analyzed to exclude memorization or overlap effects.
- The control baseline is incomplete: The in-domain text control matches the additional 500 optimization steps but may differ from synthetic PPT in sequence length, token diversity, loss scale, and difficulty. Additional controls matched on these properties are needed.
- The negative effect of Set is unexplained: Set causes large downstream degradation despite being a structured synthetic task, but the paper does not determine whether this results from harmful token-frequency statistics, excessive repetition, objective mismatch, catastrophic interference, or poor optimization.
- The relationship between PPT and data ordering is unexplored: The study does not test whether placing synthetic PPT after an initial natural-language phase, interleaving PPT with PT, or using multiple PPT stages changes the benefit.
- Retention of PPT-induced representations is not measured: The paper does not track parameter, activation, attention, or circuit changes during PT to determine whether retrieval-related structures persist, are rewritten, or are repeatedly re-learned.
- The practical applicability to modern multi-stage training is incomplete: Only broad first-stage mixtures are studied, whereas real systems often use later stages with curated data, annealing, instruction data, or domain-specific continuation. PPT’s interaction with these stages remains open.
- The conclusions may not extend beyond English web-centered training: All main evaluations and most PT data are English-oriented, leaving cross-lingual transfer and the role of language-specific grammatical or retrieval structures unresolved.
Practical Applications
Immediate Applications
The paper supports the following applications that can be implemented with existing language-model training infrastructure, although each should be validated on the target model, corpus, and evaluation tasks.
- More compute-efficient pretraining for foundation models — AI infrastructure and software
- Add a short synthetic pre-pretraining phase before conventional language-model pretraining, using tasks such as
-Shuffle Dyck, MP-Struct Core, or neural cellular automata. - The reported procedure uses approximately 500 synthetic-training steps, resets the optimizer state, and then begins ordinary pretraining. At the 3B scale, the approach reportedly achieves comparable performance while saving at least 21B subsequent pretraining tokens.
- Potential product/workflow: a pretraining configuration in frameworks such as Nanotron, Megatron-LM, or similar systems with a
synthetic_warmup_datasetoption. - Assumptions/dependencies: the synthetic task must induce specific-position, long-range retrieval rather than merely duplicate-token tracking; the implementation must preserve comparable optimizer, scheduler, tokenizer, and data-mixture settings. Savings may vary by architecture, seed, learning-rate schedule, and target corpus.
- Add a short synthetic pre-pretraining phase before conventional language-model pretraining, using tasks such as
- Low-cost improvement of open-weight LLMs — software and model release pipelines
- Open-model developers can insert retrieval-oriented PPT into existing training recipes for models in the 500M–7B range without redesigning the architecture or tokenizer.
- The reported gains persist at 7B parameters and after training budgets of up to 100B tokens, suggesting that PPT is not restricted to small, undertrained models.
- Potential product: an open-source “PPT checkpoint” or reusable initialization library distributed through Hugging Face, allowing multiple downstream models to start from a retrieval-enhanced initialization.
- Assumptions/dependencies: benefits should be measured against a same-scale, same-mixture PT-only baseline; the study does not establish that every architecture or tokenizer benefits equally.
- Improved long-context retrieval in language-model applications — search, assistants, and document QA
- Deploy PPT-trained models in workflows that require locating and reproducing information from earlier context, such as document question answering, retrieval-augmented generation, meeting summarization, contract analysis, and long-context chat.
- The strongest reported relationship is with tasks requiring information from preceding passages, including ReCoRD, HellaSwag, LAMBADA, and verbatim retrieval.
- Potential tools: retrieval-focused model evaluation suites, context-recall dashboards, and model-selection criteria based on long-range retrieval rather than only perplexity or grammatical acceptability.
- Assumptions/dependencies: improved synthetic-sequence retrieval must transfer to the particular context length, domain, language, and retrieval mechanism used in deployment. PPT does not guarantee factual accuracy, resistance to distraction, or improved external search retrieval.
- More effective training for code-and-math-capable models — developer tools and scientific AI
- Apply PPT before mixed web, code, and mathematics pretraining. The paper finds that increasing mathematics from 1.3% to 17% did not remove the PPT benefit, indicating that code and math data do not automatically make synthetic initialization redundant.
- Potential products: coding assistants, theorem-oriented LLMs, scientific literature assistants, and models trained on technical documentation.
- Assumptions/dependencies: the presence of web text appears important; removing web data largely collapsed the gain. Therefore, PPT should not be assumed to work equally well for models trained exclusively on code, mathematics, or other non-web domains.
- Training-data and initialization ablation workflow — academic and industrial research
- Use the paper’s comparison among PT-Only, ordinary text warm-up, and synthetic PPT to distinguish gains caused by extra optimization from gains caused by the synthetic task itself.
- This can become a standard experiment in model development:
- 1. train a PT-only baseline;
- 2. train a same-budget natural-text control;
- 3. train one or more retrieval-oriented PPT variants;
- 4. compare downstream, long-range retrieval, and stability metrics.
- Potential tool: an automated ablation harness reporting downstream score, token-equivalent savings, checkpoint variance, BLiMP-style linguistic scores, and retrieval NLL.
- Assumptions/dependencies: reliable conclusions require multiple random seeds and evaluation across training checkpoints. The paper used single runs for most configurations, so production decisions should not rely on one experiment.
- Evaluation of LLMs for retrieval-sensitive tasks — academia and model auditing
- Replace the assumption that grammatical acceptability is a sufficient indicator of useful structural transfer with direct tests of long-range retrieval.
- Researchers and model auditors can include synthetic span-retrieval tests, position-specific copying tasks, document QA, and distractor-context evaluations in model cards and procurement benchmarks.
- Assumptions/dependencies: retrieval benchmarks should vary distance, interference, token type, and context length; otherwise they may measure memorization or short-range pattern matching rather than general retrieval.
- Improved everyday AI assistants for long documents — daily life
- PPT-initialized models could support consumer applications such as summarizing lengthy reports, finding details in personal notes, answering questions about manuals, and recovering earlier instructions in extended conversations.
- Assumptions/dependencies: these are indirect applications based on benchmark results, not direct user studies. Privacy, hallucination control, document permissions, and domain adaptation remain necessary.
Long-Term Applications
The following applications require additional research, scaling, product engineering, or evidence beyond the paper’s experiments.
- Retrieval-specialized foundation-model pretraining curricula — AI research and infrastructure
- Develop curricula that progressively increase retrieval distance, distractor density, nesting, and sequence complexity instead of relying on a single synthetic formal language.
- Candidate tasks could include:
- keyed retrieval from earlier positions;
- multi-hop symbolic state reconstruction;
- long-range pointer tracking;
- structured sequence continuation with controlled interference;
- synthetic document and database traversal.
- Potential outcome: a general-purpose initialization stage optimized specifically for context retention and retrieval.
- Dependencies: researchers must determine which synthetic properties transfer across languages, modalities, architectures, and context lengths, and whether the benefit comes from retrieval circuits, attention allocation, optimization geometry, or another mechanism.
- Long-context models for enterprise document intelligence — legal, finance, healthcare, and public administration
- Combine retrieval-oriented PPT with long-context architectures, external retrieval systems, and domain-specific continued pretraining for use cases such as:
- legal clause and precedent retrieval;
- financial filing analysis;
- clinical-record summarization;
- regulatory compliance monitoring;
- government archive search.
- Potential product: a document-intelligence model that reports both an answer and the exact earlier passages used to produce it.
- Dependencies: high-stakes deployment requires citation fidelity, privacy controls, domain validation, robustness to adversarial documents, and regulatory compliance. The paper does not test specialized or sensitive domains.
- Adaptive training-budget allocation and carbon reduction — AI infrastructure and policy
- If token savings generalize, training providers could use PPT to reach a target quality with fewer web, code, and mathematics tokens, reducing GPU time, energy use, and data-transfer costs.
- Potential workflow: estimate the token-equivalent gain from PPT during early pilot runs, then reduce later pretraining duration while preserving a target benchmark score.
- Dependencies: the reported “21B-token saving” is specific to a 3B model and a particular comparison. Real savings must account for synthetic-data generation, validation, failed runs, checkpoint storage, and the cost of evaluating alternative PPT tasks.
- PPT-aware model pricing and compute procurement — industry and policy
- Cloud providers and research institutions could offer training recipes or compute estimates that incorporate synthetic initialization as a standard cost-saving option.
- Public funding programs could encourage reporting of token-equivalent efficiency, energy per benchmark point, and compute savings rather than only final model quality.
- Dependencies: standardized accounting is needed. Token counts alone do not capture hardware utilization, data preprocessing, memory demands, or the environmental cost of generating synthetic data.
- Multilingual and cross-domain retrieval initialization — education, translation, and global services
- Generate synthetic tasks with language-neutral symbols or multilingual structural cues, then test whether retrieval improvements transfer to low-resource languages, multilingual models, and cross-lingual document search.
- Potential products: multilingual educational tutors, translation systems that preserve earlier context, and cross-language research assistants.
- Dependencies: the present study uses one tokenizer and primarily evaluates English-oriented benchmarks. Transfer may be affected by tokenization, morphology, script, word order, and the quality of multilingual pretraining data.
- Retrieval-enhanced embodied and multimodal models — robotics and autonomous systems
- Adapt the principle to sequences of sensor states, actions, object identities, or environment transitions. A synthetic warm-up task could require retrieving a specific earlier state or matching an action to its originating observation.
- Potential applications: robot control, navigation, video understanding, event-memory systems, and multimodal agents that must recall earlier visual or proprioceptive information.
- Dependencies: the paper studies autoregressive text models only. It remains unknown whether the effect transfers to vision-language, speech, action, or world-model architectures, and whether synthetic state sequences resemble the temporal statistics of real environments.
- Reliable agent memory and tool-use systems — software automation and robotics
- Train agents on synthetic interaction traces that require locating earlier tool outputs, user constraints, permissions, or intermediate plans.
- This could improve assistants that maintain state across long workflows, such as software debugging, data analysis, or multi-step administrative tasks.
- Dependencies: better retrieval does not ensure correct planning or safe action. Systems would need explicit memory provenance, conflict resolution, access control, and tests against stale or malicious context.
- Educational systems that retain student history — education
- Use retrieval-oriented initialization for tutors that must recall earlier explanations, misconceptions, assignments, and learning goals across long sessions.
- Potential product: a tutoring model that retrieves the relevant prior interaction before generating a personalized explanation.
- Dependencies: longitudinal student data introduces privacy and consent requirements. The paper does not show improved pedagogy, personalization, or factual correctness, so educational effectiveness would require controlled user studies.
- A revised theory of synthetic pre-pretraining — academia
- Investigate whether PPT primarily changes attention patterns, positional representations, optimization trajectories, or long-range information-routing circuits rather than inducing a grammatical prior.
- Future work could use mechanistic interpretability, causal ablations of attention heads, controlled context-length experiments, and transfer tests across unrelated modalities.
- Dependencies: current evidence is correlational: retrieval improvements coincide with downstream gains, but the causal mechanism is not fully established. The weak and sometimes negative BLiMP results also indicate that grammatical competence and retrieval should be treated as separate capabilities.
- Robustness and safety evaluation for synthetic initialization — policy and model governance
- Establish standards requiring developers to test whether PPT changes:
- memorization of training sequences;
- privacy leakage from long contexts;
- susceptibility to prompt injection;
- retrieval of obsolete or conflicting instructions;
- behavior under very long adversarial contexts.
- Potential policy tool: a model-card section documenting synthetic pretraining tasks, token-equivalent savings, retrieval performance, and failure modes.
- Dependencies: improved retrieval could amplify both useful and harmful information recall. Safety implications cannot be inferred from the paper’s benchmark improvements alone.
Glossary
- Autoregressive LLM: A LLM that predicts each token from the tokens preceding it. “Let be an autoregressive LM with weights ”
- Batch size: The number of training examples processed in one optimization step. “we raise the batch size to 512 sequences”
- BLiMP: A benchmark that evaluates whether LLMs distinguish grammatically acceptable from unacceptable sentences. “BLiMP measures zero-shot grammatical acceptability over 12 paradigm groups”
- Context-sensitive language: A formal language whose valid strings may require context-dependent rules to recognize. “-Shuffle Dyck is a context-sensitive language of interleaved bracket pairs”
- Constituency: The hierarchical organization of words into syntactic phrases. “These are the exact dependencies that -Shuffle Dyck, a hierarchical bracket-matching language, is hypothesized to transfer.”
- Context window: The maximum sequence of tokens that a model can process simultaneously. “with a 4,096-token context window”
- Curriculum: A training strategy in which data or tasks are presented in successive stages, often with changing difficulty or composition. “Modern PT approaches use a multi-stage curriculum”
- Decoder-only LLM: A Transformer LLM that generates text autoregressively using only decoder-style layers. “In modern decoder-only LMs, PPT is a warm-up training phase”
- Downstream task: A task used to measure the capabilities obtained after a model has been pretrained. “We evaluate models across ten downstream benchmarks”
- Embedding matrix: A trainable matrix that maps discrete tokens to continuous vector representations. “with the embedding matrix and LM head re-initialized for the natural language tokenizer”
- Formal grammar: A precisely specified system of rules defining which sequences belong to a language. “two based on formal languages, two structured synthetic tasks without formal grammars”
- Formal language: A language defined by symbolic rules rather than by natural human communication. “typically from a formal language such as -Shuffle Dyck”
- Gradient descent: An optimization method that updates model parameters in the direction that reduces a loss function. “large Transformers can learn hierarchical abstractions directly through standard gradient descent”
- Gradient noise: Random variation in parameter updates caused by estimating gradients from finite batches of data. “introducing higher stochastic gradient noise than large-batch setups”
- Grammatical acceptability: The degree to which a sentence conforms to the grammatical rules perceived by speakers or evaluated by a benchmark. “measures zero-shot grammatical acceptability over 12 paradigm groups”
- Grammatical prior: A pre-existing structural bias toward grammatical patterns that is hypothesized to be learned during synthetic pre-pretraining. “PPT is claimed to induce a grammatical prior that transfers to natural language.”
- Hierarchical abstraction: A representation of relationships organized across multiple nested structural levels. “large Transformers can learn hierarchical abstractions directly through standard gradient descent”
- Hierarchical dependency: A relationship between elements whose interpretation depends on nested or multi-level structure. “an effective corpus should capture hierarchical dependencies”
- Inductive bias: A preference or structural assumption that guides a model toward particular types of solutions. “a structural inductive bias acquired from PPT data”
- Initialization state: The parameter values from which model training begins. “These weights then replace as the initialization state for PT.”
- Language modeling: The task of assigning probabilities to sequences of tokens, typically by predicting each token from prior context. “Language Modeling (LM): LAMBADA”
- Long-range retrieval: The ability to locate and reproduce information appearing substantially earlier in a sequence. “This suggests that PPT induces a long-range retrieval capability”
- Loss function: A numerical objective measuring how poorly a model’s predictions match its training data. “Standard PT minimizes $\mathcal{L}(\theta; \mathcal{D}_{\text{PT})$”
- Negative log-likelihood (NLL): A loss obtained by taking the negative logarithm of the probability assigned to the observed data. “denote the negative log-likelihood (NLL) over a corpus ”
- Neural cellular automata (NCA): Systems composed of cells whose states are repeatedly updated according to local neural rules. “The neural cellular automata (NCA) task consists of successive states of an NCA”
- Parameter scaling: Increasing the number of trainable parameters in a model to study how performance changes with model size. “the downstream performance and token efficiency benefits of PPT persist under parameter scaling”
- Pre-pretraining (PPT): An initial training phase on synthetic or auxiliary data performed before standard language-model pretraining. “Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency”
- Pre-training (PT): The main training process in which a LLM learns from a large corpus before evaluation or adaptation. “Standard PT minimizes $\mathcal{L}(\theta; \mathcal{D}_{\text{PT})$”
- Random initialization: Starting model training with parameters assigned randomly rather than inherited from another trained model. “Each PPT run optimizes $\mathcal{L}(\theta; \mathcal{D}_{\text{PPT})$ from a random initialization ”
- Retrieval ambiguity: Uncertainty about which earlier item a current item should be matched with or associated with. “reducing retrieval ambiguity in those dependencies is what drives the gain”
- Semantic subgroup: A category of linguistic evaluation concerned with meaning and interpretation. “reporting overall mean accuracy alongside subgroup scores for semantics, morphology, and syntax”
- Stochastic gradient noise: Variability in estimated gradients caused by random sampling during mini-batch optimization. “introducing higher stochastic gradient noise than large-batch setups”
- Structural abstraction: A generalized representation of relationships or patterns underlying particular data instances. “Such PT mixtures may induce structural abstractions”
- Structural inductive bias: A model tendency to favor particular forms of organization or dependency structure. “a structural inductive bias learned during PPT that transfers to natural language grammar”
- Synthetic corpus: A dataset generated artificially according to specified rules rather than collected from natural language. “a generator produces a synthetic corpus”
- Token efficiency: The amount of useful performance gained per training token, or the number of tokens needed to reach a given performance level. “improves token efficiency during LLM pre-training”
- Tokenizer: A system that converts text into discrete tokens used as model inputs. “All models use the SmolLM3 architecture and tokenizer”
- Verbatim retrieval: The task of reproducing an earlier sequence exactly rather than merely recalling its meaning. “Verbatim retrieval instead requires locating a span earlier in a long context and copying it.”
- Warm-up phase: An initial training period used to prepare model parameters before the main training procedure. “PPT is a warm-up training phase on synthetic sequences before PT”
- Zero-shot evaluation: Testing a model on a task without providing task-specific examples during evaluation. “BLiMP measures zero-shot grammatical acceptability over 12 paradigm groups”

