REER-PT: Reverse-Engineered Reasoning for Perplexity-Guided Pre-training Data Augmentation
Abstract: As language-model compute continues to scale, high-quality training data is becoming an increasingly important bottleneck. Conventional next-token prediction supervises what follows a context but leaves the intermediate reasoning behind that continuation implicit. We introduce \textbf{REER-PT}, a scalable framework that extends Reverse-Engineered Reasoning (REER) to raw pre-training data. REER-PT identifies continuations that are difficult to predict but can still be inferred from the preceding context, and inserts concise reasoning annotations that reconstruct the missing connection between context and continuation. Candidate annotations are generated and refined offline, with perplexity serving as the optimization signal. Constraints on length and target leakage filter out unhelpful or trivial annotations. This sparse transformation preserves the source text and remains compatible with standard next-token prediction, avoiding online reasoning rollouts during pre-training. We apply REER-PT to transform a source pre-training corpus into an augmented one. Across augmented-data, original-token, and selected-continuation comparisons, perplexity reductions range from 0.42 to 7.29, and only about 0.05\% of annotation 13-grams appear verbatim in the source text. We then train two 680M-parameter models with the same architecture and training configuration on the source and augmented corpora, respectively. The augmented-data model gains up to 2.07 percentage points on several knowledge and reasoning benchmarks. Together, the perplexity analysis indicates improved continuation predictability, while the controlled pre-training experiments suggest that this augmentation can improve model performance without changing the standard pre-training objective.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces a method called REER-PT for improving the data used to train LLMs.
LLMs learn by reading huge amounts of text and trying to predict the next word. For example:
“The sky became dark, so…”
A model might learn that “it started raining” often comes next. However, the training text usually does not explain why the next sentence makes sense.
REER-PT adds short explanations between parts of existing text. These explanations act like helpful notes that show the connection between what came before and what comes next. The original text is not deleted or rewritten.
2. What questions are the researchers asking?
The researchers mainly want to know:
- Can short reasoning notes make difficult parts of training text easier for a LLM to understand?
- Can these notes be added to very large collections of ordinary documents, rather than only to special question-and-answer datasets?
- Can the notes improve the LLM’s knowledge and reasoning abilities?
- Can this be done efficiently, without making the model “think out loud” during training?
- Do the added notes accidentally copy too much of the original text?
- Does this method work equally well for natural language and computer code?
The basic idea is similar to adding a teacher’s note to a textbook. The note explains an important connection, but the original textbook sentences remain unchanged.
3. How did the researchers do it?
Finding places that need help
First, the researchers used a LLM to read documents and measure how difficult each sentence was to predict.
They used a measurement called perplexity. In simple terms, perplexity tells us how surprised a LLM is by a piece of text:
- Low perplexity: The next words are fairly easy to guess.
- High perplexity: The next words are difficult or surprising to guess.
However, a surprising sentence is not always a good place for an explanation. For example, a sentence containing a person’s unusual name or a random identification number may be hard to predict even when no explanation could help.
Therefore, the researchers kept only sentences that were:
- Difficult for the model to predict, and
- Still understandable from the earlier context.
Creating explanation notes
For each selected place, another model wrote several possible explanations. These were designed to sound like short, neutral notes in a book.
For example, imagine that a text first discusses insulin and glucagon, and then suddenly talks about how the liver controls sugar levels. The added note might explain that insulin and glucagon affect liver activity, which connects the earlier discussion to the next part.
The notes were supposed to:
- Explain the missing connection,
- Be concise,
- Avoid simply repeating the next sentence,
- Avoid revealing the continuation directly.
The paper calls directly giving away the next part target leakage. This would be like writing the answer to a puzzle in the hint instead of helping someone understand how to solve it.
Testing and improving the notes
The researchers then tested each possible note by asking:
Does this note make the following text easier for the LLM to predict?
If the note lowered the continuation’s perplexity, it was considered useful. The researchers also repeatedly rewrote parts of the note and kept versions that worked better.
This search happened offline, before training the final LLM. That means the model did not need to generate explanations while it was being trained.
Building the new training dataset
The researchers applied this process to about 23 billion tokens of original text. A token is a small piece of text, such as a word, part of a word, or punctuation mark.
After adding the notes, the dataset grew to about 42 billion tokens. The original text stayed in the same order, with the new explanations inserted before selected passages.
Finally, they trained two LLMs:
- One model learned from the original dataset.
- The other learned from the dataset with REER-PT notes.
Both models had about 680 million parameters and used the same training setup. This made the comparison fairer.
4. What did they find?
The added notes made text easier to predict
The researchers found that the notes lowered perplexity in several ways.
For the full augmented dataset, perplexity dropped by about 7.29 points compared with the original text. When looking only at the original text—not the added notes—perplexity still dropped by about 1.03 points.
This is important because it suggests that the notes did not merely make the newly added text easy to read. They also helped the model predict the unchanged original text.
For the selected difficult passages, perplexity dropped by about 4.24 points. This means the notes were especially useful at the places where the model had trouble understanding how one idea led to the next.
The notes did not copy much of the source text
The researchers checked groups of 13 words or characters, called 13-grams, to look for repetition.
They found that only about 0.051% of the note text exactly matched a sequence from the original document. This is a very small amount, suggesting that the notes usually explained ideas in new words instead of copying the source.
The improved model performed better on knowledge and reasoning
The model trained with REER-PT data performed better on many tests involving knowledge, reasoning, mathematics, and science.
Some of the largest improvements were:
| Test | Improvement |
|---|---|
| BBH, a general reasoning test | +2.07 points |
| GPQA-Diamond, a difficult science test | +2.07 points |
| MATH | +1.50 points |
| OlympiadBench | +1.49 points |
| DROP, a reading and reasoning test | +1.40 points |
| MMLU-Pro, a broad knowledge test | +0.90 points |
These results suggest that helping a model understand connections between ideas during training may improve its ability to answer questions and solve problems later.
Code performance became worse
The method did not improve every ability. The model performed worse on all three programming tests:
- MBPP+: −2.65 points
- HumanEval+: −1.83 points
- LiveCodeBench: −1.79 points
The researchers think this happened because natural-language notes were inserted into code documents. Those notes may confuse the model about where the program begins and ends. As a result, the model might produce explanations mixed together with code, causing the code to fail even when the general idea is correct.
5. Why is this research important?
LLMs need enormous amounts of high-quality training data. Creating specially written reasoning examples by hand would take too much time and effort.
REER-PT offers a way to improve ordinary documents automatically. Its advantages include:
- It can work on very large datasets.
- It keeps the original text.
- It adds explanations only where they seem useful.
- It uses the normal training method of predicting the next token.
- It does not require expensive reasoning during the model’s training process.
- It may improve knowledge, mathematics, science, and general reasoning abilities.
In everyday terms, REER-PT is like placing helpful study notes in difficult parts of a giant textbook collection. The notes show how ideas connect, making the material easier for an AI student to learn from.
However, the results should be treated carefully. The researchers tested only one model size, one main training setup, and a limited type of annotation. The method also hurt code-generation performance, so code may need special notes designed to preserve programming structure.
Conclusion
The paper argues that LLMs can learn better when training text includes short explanations of hidden connections between ideas. REER-PT finds difficult but understandable passages, creates possible notes, and keeps only the notes that help the model predict what comes next.
The method improved several knowledge and reasoning tests, with gains of up to 2.07 percentage points. Its biggest weakness was programming performance, because ordinary explanatory notes can interfere with code.
Overall, this research suggests that improving the quality and structure of training data—not just adding more data—could help build more capable LLMs.
Knowledge Gaps
Knowledge Gaps, Limitations, and Open Questions
- Scaling behavior is unresolved: The method is evaluated only with 680M-parameter models and a single training recipe; its effectiveness, cost, and stability for multi-billion-parameter models remain unknown.
- Limited replication evidence: The comparison relies on one raw model and one augmented-data model, with no multiple random seeds or statistical significance analysis to establish whether benchmark gains are robust.
- Confounded training comparisons: The augmented mixture contains approximately 19B more tokens than the raw mixture, so improvements cannot be attributed solely to reasoning annotations rather than increased token exposure or compute.
- No compute- or token-matched baseline: The paper does not compare REER-PT against a baseline trained for the same number of optimization steps, consumed tokens, or FLOPs, nor against a corpus with equally sized non-reasoning augmentations.
- Contribution of each pipeline component is unclear: There is no full ablation isolating sentence-level perplexity selection, inferability filtering, annotation generation, leakage filtering, refinement, and final perplexity-based acceptance.
- The choice of selection budget is not justified: The fixed rule and the resulting annotation density are not compared with alternative densities or adaptive budgets.
- Annotation length effects are unexplored: The typical 500–1,000-word constraint is not systematically varied, leaving unclear whether shorter explanations, longer reasoning traces, or token-budget-matched alternatives are more effective.
- Dependence on the PPL model is undercharacterized: Although the paper notes that different PPL models may select different positions, it does not quantify the sensitivity of selected data, annotations, or downstream performance to model size, domain, checkpoint, or on-policy/off-policy status.
- Reward-model overfitting is possible: Refinement directly optimizes continuation perplexity under one PPL model, but the paper does not test whether gains transfer to other LLMs or merely exploit quirks of the evaluator.
- Perplexity reduction is not validated as causal evidence of reasoning quality: Lower continuation perplexity may result from stylistic priming, discourse regularization, or hidden target cues rather than genuinely recovering an intermediate dependency.
- Inferability judgments are insufficiently validated: The paper does not report human agreement, evaluator accuracy, or error analysis for determining whether a continuation is actually inferable from its context.
- Target-leakage detection is underspecified: Leakage is described as repetition or close paraphrasing, but the detection procedure, thresholds, semantic evaluation, and residual leakage rates are not reported in enough detail to establish that annotations do not reveal the continuation.
- Exact 13-gram overlap is an incomplete contamination measure: The low 13-gram overlap does not rule out shorter phrase copying, semantic paraphrase, fact copying, or leakage through names, numbers, and structured information.
- Annotation factuality and faithfulness are not evaluated: Generated annotations may introduce unsupported claims, distort the source context, or provide plausible but incorrect explanations; no systematic factuality or attribution assessment is presented.
- Quality variation across domains is unexplored: The corpus-level results do not show how annotation quality and utility differ across scientific, educational, news, legal, conversational, multilingual, or low-resource documents.
- The source-corpus composition is insufficiently described: Details about data sources, language distribution, filtering, duplication, and domain proportions are missing, limiting reproducibility and assessment of external validity.
- Long-context effects are unknown: The study does not examine whether inserting 500–1,000-word annotations causes context-window truncation, changes document-level dependencies, or harms information retrieval over long documents.
- Interactions between multiple annotations are not isolated: Because several annotations may be inserted into one document, it remains unclear whether gains arise from local bridging, cumulative document restructuring, or interactions among annotations.
- Potential distribution shift from artificial markers is unexamined: The effects of
<annotation_begin>and<annotation_end>on pre-training behavior, downstream prompting, generation style, and deployment-time outputs are not measured. - The method’s impact on generation quality is incomplete: Evaluation focuses mainly on benchmark accuracy and code-generation scores; open-ended fluency, factuality, calibration, instruction following, toxicity, verbosity, and stylistic contamination are not assessed.
- Code degradation lacks a controlled diagnosis: The reported regressions on code benchmarks are attributed to natural-language insertions, but there is no comparison with code-specific annotations, comments, separate channels, or structure-preserving placement strategies.
- Effects on other structured domains remain unknown: The paper does not test whether the same approach harms or benefits tables, formulas, markup, mathematical notation, dialogue, or multimodal-associated text.
- No comparison with competing data-transformation methods is provided: REER-PT is not directly compared under matched budgets with rephrasing, rewriting, synthetic instructions, textbook transformations, latent reasoning, or reinforcement-pretraining approaches.
- Annotation-generation cost is not reported: The computational cost, number of candidate generations, refinement evaluations, storage overhead, and wall-clock time for transforming 23B tokens are absent, making the claimed scalability difficult to assess.
- Inference-time and training-time trade-offs are unclear: The paper does not determine whether the added annotation tokens improve capability per FLOP, capability per source token, or capability per total training token.
- The persistence of gains after fine-tuning is unknown: It is unclear whether REER-PT benefits survive supervised fine-tuning, instruction tuning, preference optimization, or reinforcement learning.
- Data freshness and temporal generalization are not studied: The method may amplify outdated or incorrect source transitions, but there is no evaluation on temporally held-out knowledge or changing factual domains.
- Robustness to noisy or adversarial source text is unresolved: The pipeline may generate and optimize misleading annotations for corrupted, contradictory, low-quality, or adversarial documents.
- The optimal annotation style is not established: The preference for third-person, book-note-style annotations is motivated qualitatively, but alternative styles and their effects on learning, generation, and domain compatibility are not systematically compared.
- Mechanistic evidence for capability gains is missing: The paper does not identify whether annotations improve factual recall, discourse modeling, multi-step reasoning, representation quality, or merely benchmark-specific pattern recognition.
- Transfer beyond the evaluated benchmarks is unknown: The benchmark set does not establish performance on broader language understanding, multilingual reasoning, long-context tasks, retrieval, safety, or real-world applications.
- Reproducibility is limited by missing implementation details: Prompts, model identities and checkpoints, candidate counts, refinement schedules, stopping criteria, filtering algorithms, training schedules, and evaluation protocols are not fully specified.
- The relationship between source difficulty and benefit is unresolved: The paper does not characterize which types of high-perplexity, contextually inferable transitions yield the largest downstream improvements or whether selection can be made more targeted.
- Failure modes of accepted annotations are not analyzed: No qualitative taxonomy is provided for annotations that reduce PPL but are redundant, misleading, overly generic, stylistically unnatural, or harmful to later generation.
- The method’s effect on memorization and privacy is unknown: Synthetic annotations may restate sensitive source information in new forms, potentially affecting privacy leakage, memorization, or copyright exposure despite low exact n-gram overlap.
Practical Applications
Immediate Applications
- Pre-training data augmentation for language-model developers (AI/software; deployable now) Organizations can integrate REER-PT as an offline preprocessing stage: calculate sentence-level perplexity, filter for contextually inferable but difficult transitions, generate candidate annotations, refine them using continuation perplexity, and insert accepted annotations between dedicated boundary markers. The resulting corpus can be trained with the existing next-token-prediction pipeline, without modifying the model architecture or adding online reasoning rollouts. Potential tools/workflows: a batch data-processing service, annotation-generation and refinement jobs, perplexity-ranking dashboards, and augmented-corpus versioning. Evidence: the paper reports reductions in perplexity and gains of up to 2.07 percentage points on several knowledge and reasoning benchmarks. Dependencies: substantial offline inference and storage costs; access to a suitable perplexity model and annotation model; reliable sentence segmentation; quality controls for hallucinations, target leakage, and annotation length.
- Improving knowledge- and reasoning-oriented foundation models (AI research and commercial NLP; deployable now for experimentation) Model developers can apply the method selectively to general prose, educational material, scientific writing, legal documents, manuals, and other expository sources where implicit transitions are common. This may improve downstream systems for question answering, summarization, document analysis, and reasoning. Potential products: domain-adapted LLMs, retrieval-augmented generation backbones, enterprise copilots, and knowledge-management assistants. Dependencies: the reported experiments use only 680M-parameter models and one training recipe, so benefits at larger scales or in production models are not established. Domain-specific factual validation and benchmark testing are required.
- Continual pre-training and domain adaptation (healthcare, finance, law, education, and enterprise software; deployable now) REER-PT can be applied to a smaller domain corpus before continual pre-training. For example, a healthcare organization could augment clinical textbooks or guidelines, while a financial institution could augment policy manuals and market research. The inserted notes could make domain-specific causal, procedural, or explanatory links more explicit for the adapted model. Potential workflow: select high-perplexity passages from a licensed corpus, have subject-matter experts review a sample of annotations, retain only annotations that pass perplexity and factuality checks, then conduct domain-adaptive pre-training. Dependencies: licensing, privacy protection, expert review, and strict prevention of unsupported medical, legal, or financial claims. Perplexity reduction alone does not establish factual correctness.
- Corpus quality auditing and prioritization (data engineering and academic infrastructure; deployable now) The selection stage can be used independently of annotation insertion to identify passages that are difficult for a model but potentially explainable from their context. Data teams can use these scores to prioritize human review, deduplication, rewriting, or source-quality investigations. Potential tools: “reasoning opportunity” maps, document-quality queues, and token-level curriculum schedules. Dependencies: high perplexity also occurs for arbitrary names, dates, identifiers, and genuinely external facts; therefore, inferability filtering and human or automated validation are necessary.
- Offline evaluation of explanatory annotations (academia and model governance; deployable now) Researchers can compare candidate explanations by measuring whether they lower the perplexity of an unchanged continuation while avoiding direct repetition. This provides a quantitative screening mechanism for synthetic training text and can complement factuality, attribution, and human-quality evaluations. Dependencies: the PPL model may favor stylistic or distributional artifacts rather than genuinely useful explanations. Results should be reported alongside independent evaluations, ablations, and contamination checks.
- Educational content enrichment and study tools (education; deployable now for low-risk content) Textbooks, lecture notes, and public educational material can be augmented with concise intermediate explanations that clarify why one concept, result, or paragraph follows another. These annotations could support teacher-authoring systems, reading assistants, adaptive study guides, and question-generation pipelines. Dependencies: annotations require review by educators, particularly in mathematics, science, and history. The method does not guarantee pedagogical appropriateness, age suitability, or factual accuracy.
- Document navigation and comprehension assistance (daily life and knowledge work; deployable now) A production system could use the same pipeline to generate “why this follows” notes for long reports, technical manuals, policies, or articles. Users could receive short explanations of implicit transitions before a new section, improving skimming and comprehension. Dependencies: this application is an extrapolation from pre-training results rather than a directly evaluated user study. Notes must clearly be labeled as generated, grounded in the source, and checked for sensitive or misleading interpretations.
- Avoiding natural-language augmentation in code corpora (software engineering; immediate negative application) The findings provide an actionable constraint: do not apply the standard book-note format indiscriminately to source code. The augmented model underperformed the baseline on MBPP+, HumanEval+, and LiveCodeBench, likely because inserted prose disrupted program structure and encouraged explanatory text in generated code. Practical workflow: route code documents to a separate preprocessing policy, preserve syntax and repository structure, and evaluate code-generation benchmarks independently from natural-language benchmarks.
Long-Term Applications
- Structure-aware reasoning augmentation for code and configuration files (software engineering and developer tools; requires further research) A future variant could insert comments, docstrings, type-level explanations, dependency summaries, or repository-level design notes at syntactically valid locations rather than inserting ordinary prose into code. This could support code-completion models, debugging assistants, migration tools, and software-maintenance agents. Dependencies: language-specific parsers, AST- or repository-aware placement, strict separation of executable code and explanatory content, and safeguards against changing program semantics. The current paper’s negative code results show that the standard format is unsuitable.
- Specialized augmentation formats for technical domains (healthcare, law, finance, robotics, and engineering; requires further research) Domain-specific annotations could represent causal mechanisms in medical texts, regulatory dependencies in legal documents, assumptions in financial reports, or physical constraints in engineering manuals. Such formats may be more useful than generic book-note prose. Dependencies: expert-designed schemas, domain validators, provenance tracking, and tests for whether the annotations improve real task performance rather than only perplexity.
- Scaling REER-PT to large foundation models and larger corpora (AI infrastructure; requires further research and engineering) The method could become part of large-scale data factories for trillion-token training runs, potentially using distributed inference, caching, approximate perplexity estimation, and adaptive annotation density. The target ratio of roughly one insertion per 1,000 source tokens could be adjusted by domain and model size. Dependencies: the paper does not establish scaling laws, cost-effectiveness, optimal density, or whether gains persist for larger models. Additional experiments should control for the increased token count: the raw and augmented training mixtures contained approximately 523B and 542B tokens, respectively.
- On-policy, model-specific training-data construction (foundation-model development; requires further research) Using the eventual target model as the perplexity model could select annotations tailored to that model’s actual weaknesses. Periodic checkpoints might enable iterative corpus improvement during training or continual pre-training. Dependencies: repeated data construction is expensive and may amplify model-specific biases or errors. Researchers would need to determine whether off-policy annotations transfer better across architectures and whether on-policy optimization causes overfitting.
- Reasoning-aware curricula and adaptive data mixtures (academia and AI training systems; requires further research) Perplexity and inferability scores could drive a curriculum that gradually introduces difficult but explainable transitions, while excluding unpredictable content that cannot be reconstructed from context. This could combine REER-PT with data selection, deduplication, or excess-loss-based sampling. Dependencies: curriculum benefits are not directly tested. Poor calibration of perplexity may cause the system to over-select stylistically unusual text or under-select valuable factual material.
- Provenance-preserving synthetic-data standards (policy, research governance, and enterprise compliance; requires further development) Augmented corpora could retain the original document, annotation text, insertion location, model versions, perplexity scores, filtering decisions, and source licenses. Such metadata would support audits of synthetic data, copyright review, dataset documentation, and reproducible model training. Dependencies: exact 13-gram overlap is a limited measure of copying and does not resolve copyright, attribution, or factuality concerns. Governance frameworks would need stronger semantic similarity, provenance, consent, and risk assessments.
- Human-in-the-loop scientific and policy knowledge systems (academia and public policy; requires further research) Expert reviewers could inspect high-value context-to-continuation transitions and approve annotations for policy documents, scientific literature, standards, or public information services. Approved annotations could improve models used for evidence synthesis and cross-document reasoning. Dependencies: expert review costs may be substantial; annotations must preserve uncertainty, distinguish evidence from inference, and avoid presenting generated connections as authoritative conclusions.
- Personalized reading and accessibility systems (daily life, accessibility, and lifelong learning; requires product development) A future reading assistant could generate explanations at different complexity levels, identify omitted discourse links, and adapt notes to a reader’s background. This may benefit users reading technical, multilingual, or cognitively demanding material. Dependencies: user studies are needed to verify comprehension gains. Systems must avoid over-explaining, hallucinating connections, exposing private documents, or replacing professional interpretation in high-stakes domains.
Glossary
- 13-gram: A sequence of 13 consecutive words, characters, or tokens used to measure repetition and overlap. “We use 13-grams throughout.”
- Annotation model: A LLM responsible for checking inferability and generating or revising reasoning annotations. “The annotation model performs the inferability check and generates both initial annotations and candidate rewrites during refinement.”
- Augmented corpus: A training corpus created by inserting additional annotations into the original text. “The resulting augmented corpus is fixed before pre-training and can be used directly with the standard next-token prediction objective.”
- Candidate annotation space: The set of possible annotations considered for a particular context–continuation pair. “Let denote the candidate annotation space for the -th context--continuation pair.”
- Chain-of-thought (CoT) supervision: Training supervision that provides explicit intermediate reasoning steps leading to an answer or continuation. “Chain-of-thought (CoT) supervision provides a natural mechanism for representing these dependencies explicitly.”
- Conditional next-token distribution: The probability distribution assigned to the next token given the preceding token sequence. “We denote its conditional next-token distribution by .”
- Contextual inferability: The degree to which a continuation can be logically or semantically inferred from its preceding context. “This stage selects insertion positions that are both difficult to predict and contextually inferable.”
- Continuation perplexity: The perplexity of a target continuation conditioned on its preceding context and, optionally, an annotation. “Following REER \citep{wang2025reer}, we use the perplexity of the observed continuation as the optimization signal.”
- Corpus augmentation: The process of modifying a text corpus by adding information intended to improve model training. “Effective corpus augmentation must therefore determine both where reasoning is useful and whether a proposed annotation makes the observed continuation easier to predict.”
- Exact overlap: The proportion of text sequences that occur identically in two texts. “The mean annotation-to-source exact-overlap ratio is 0.051\%, indicating that only a small fraction of annotation 13-gram occurrences exactly match a source span.”
- Excess loss: The amount by which a token’s loss exceeds a reference model’s corresponding loss. “Rho-1 applies selective language modeling, using token-level excess loss relative to a reference model to concentrate optimization on high-excess-loss portions of the corpus.”
- Gradient-free search: An optimization procedure that searches for improvements without using gradients. “Enumerating all possible annotations is infeasible, so we perform an iterative, gradient-free search.”
- Gradient norm: The magnitude of the gradient vector used to update model parameters during optimization. “The left and center panels show raw training loss and gradient norm over the complete runs.”
- Inferability check: A procedure that determines whether a continuation is adequately supported by its preceding context. “The annotation model therefore performs an inferability check and filters out candidates that cannot be supported by the context.”
- Insertion position: A location in a document where an annotation may be placed before a continuation. “At the -th selected position, denotes the local context preceding the boundary, denotes the observed continuation following it, and denotes the final accepted annotation inserted between them.”
- Latent chain of thought: Intermediate reasoning representations that are not directly exposed as ordinary text. “Adaptive latent CoT \citep{zeng2026adaptivecot} allocates variable-length latent reasoning according to token difficulty.”
- Log probability: The natural logarithm of a model-assigned probability, commonly used in language-model objectives. “For each token, we compute .”
- Mid-training: A training phase occurring between initial pre-training and later post-training or alignment stages. “Recent work also brings reinforcement learning into pre-training and mid-training by deriving rewards or learning signals from pre-training corpora.”
- Next-token prediction: The objective of predicting each subsequent token from the tokens that precede it. “LLMs acquire most of their capabilities through next-token prediction on massive text corpora.”
- Off-policy: A training or data-construction procedure whose selection or optimization model differs from the target model being trained. “Using a separate PPL model yields an off-policy procedure.”
- On-policy: A procedure in which the model used to select or optimize data is the same model, or policy, being trained. “Using the target model itself as the PPL model yields on-policy selection and optimization.”
- Perplexity (PPL): An exponential measure of a probabilistic model’s average uncertainty when predicting a sequence. “A larger indicates that the sentence is more difficult to predict from its preceding text.”
- Perplexity-guided refinement: Iteratively revising annotations according to whether they reduce the perplexity of a target continuation. “It inserts a concise reasoning annotation before a selected continuation only if the annotation reduces that continuation's perplexity.”
- Policy rollout: An online generation process in which a model produces a sequence of actions or reasoning steps. “This avoids policy rollouts and reward optimization during model training.”
- Pre-training mixture: A combined collection of datasets used to train a model during pre-training. “We construct two pre-training mixtures.”
- Prefix: The sequence of tokens preceding a particular token or continuation. “where is the token-level log probability and is the prefix preceding .”
- Reasoning annotation: Added text that explicitly describes the conceptual connection between a context and its continuation. “REER-PT inserts a concise, third-person, book-note-style annotation that captures this connection without revealing the target content.”
- Reasoning trajectory: A sequence of intermediate reasoning steps used to derive or predict an output. “REER uses the perplexity of a known reference output as an optimization signal, searching over candidate CoT trajectories for reasoning that makes the reference easier to generate.”
- Reference perplexity: The perplexity assigned to a known target output by a model, used as an optimization signal. “Reference-guided methods include REER \citep{wang2025reer}, which uses reference perplexity to search for effective reasoning trajectories.”
- Reinforcement pre-training: Pre-training that incorporates reinforcement-learning objectives or rewards derived from corpus data. “In reinforcement pre-training, RPT reframes next-token prediction as a reinforcement-learning task and rewards reasoning that correctly predicts the next token.”
- Self-repetition ratio: The fraction of n-gram occurrences that repeat beyond the first occurrence within the same text. “For any text , we define the self-repetition ratio as”
- Selective language modeling: A training strategy that gives greater emphasis to selected tokens or corpus segments. “Rho-1 applies selective language modeling, using token-level excess loss relative to a reference model to concentrate optimization on high-excess-loss portions of the corpus.”
- Target leakage: Inclusion of information from the target continuation in an annotation, allowing the target to be predicted by copying rather than reasoning. “Target leakage occurs when an annotation directly repeats or closely paraphrases words or facts from the continuation.”
- Token-level log probability: The logarithm of the probability assigned to an individual token conditioned on its preceding sequence. “where is the token-level log probability and is the prefix preceding .”
- Tokenization: The process of converting text into a sequence of model-processing units called tokens. “Let denote the token sequence obtained by tokenizing a document.”
- Verbatim copying: Reproducing text exactly rather than generating an equivalent explanation or paraphrase. “These measurements show little repetition or verbatim copying.”



