- The paper introduces token-level noise, sequence permutations, and target offset prediction as augmentations to regularize AR pretraining in data-constrained regimes.
- It demonstrates that tuning parameters, such as reducing token noise to 5%, significantly lowers validation loss and boosts downstream zero-shot performance.
- Empirical results reveal that combining right-to-left permutation with offset prediction can synergistically overcome overfitting in multi-epoch settings.
Data Augmentations for Data-Constrained LLM Pretraining: A Technical Essay
Motivation and Problem Setting
Recent advances in LLMs have been driven by escalating compute budgets and steady expansion of large-scale internet text corpora. This paradigm is running up against a hard data wall: projections indicate that high-quality web text, a primary driver for LM scaling, will be exhausted before compute availability plateaus. As GPU capacity outpaces data generation, practitioners must adapt to a compute-abundant, data-constrained regime, characterized by multi-epoch pretraining far below the Chinchilla-optimal data-to-parameter ratio. In this context, standard autoregressive (AR) pretraining exhibits severe overfitting: after a brief period of generalization, the model rapidly memorizes and validation loss increases monotonically, rendering extended training counterproductive Figure 1.

Figure 1: Validation loss of the baseline AR model over 100 epochs. The loss bottoms out at epoch~16 and deteriorates continuously thereafter.
Diffusion LMs have been shown to resist such overfitting via objective-level regularization, but are not yet practical for large-scale deployment due to generation inefficiencies and infrastructure limitations. The central question addressed is whether analogous regularization can be achieved in standard AR pretraining pipelines through systematic data augmentation, thereby allowing productive, robust multi-epoch training on fixed corpora.
Augmentation Design Space
The work introduces and analyzes three orthogonal augmentation classes, designed to decorrelate the training distribution from the raw data and thereby mitigate overfitting without altering the underlying AR loss:

Figure 2: Overview of the three augmentation categories: token-level noise, sequence permutations, and target offset prediction.
- Token-Level Noise: A subset of input tokens is corrupted at each training step, either replaced by a dedicated mask token or a random vocabulary token, but the label sequence is always the original, uncorrupted text. This mirrors the denoising paradigm in diffusion models.
- Sequence Permutations: Input sequences are either reversed entirely (R2L) or rearranged in Fill-in-the-Middle (FIM) style. Specialized control tokens signal the transformation; labels match the permuted order.
- Target Offset Prediction: The model is tasked with predicting a future token xt+iโ (rather than the immediate xt+1โ), with the offset i sampled from a fixed distribution for each sample, and the current offset signaled by a special token.
These augmentations are strictly training-time. During evaluation, no transformations are applied; all models are assessed using standard L2R next-token prediction.
Experimental Protocol
A 150M-parameter Llama-based model is trained on 75M tokens sampled from DCLM-RefinedWeb (40ร below Chinchilla-optimal data budget). All augmentations are applied within the HuggingFace Transformers implementation and evaluated with a Warmup-Stable-Decay (WSD) learning schedule. The key metric is held-out validation loss; secondary reporting involves zero-shot downstream accuracy on a suite of reasoning and commonsense benchmarks.
Analysis of Individual Augmentation Strategies
Token-Level Noise
Random token replacement outperforms masking across all corruption rates Figure 3. Notably, 15% replacement achieves a minimum validation loss of 3.841, an absolute improvement over both baseline (4.015) and masking at the same rate.

Figure 3: Validation loss for token-level noise ablations. Random replacement outperforms masking at matched rates; among random replacement variants, 15\% achieves the best individual minimum.
The effect is attributed to the increased challenge posed by plausible incorrect tokens, as opposed to the explicit information absence cues from masks. Higher rates yield diminishing returns, while rates that are too low provide insufficient regularization.
Sequence Permutations
R2L with a 50% mixing rate results in substantial regularization, reducing validation loss to 3.910, while FIM does not regularize and even exacerbates overfitting, likely due to large trainโeval distribution mismatch Figure 4.

Figure 4: Validation loss for sequence permutation ablations. R2L at 50\% provides strong regularization; FIM provides essentially no benefit and overfits at the same rate as the baseline.
Target Offset Prediction
Offset prediction is maximally effective when using an exponentially decaying weighting over iโค5, yielding a minimum loss of 3.870 Figure 5. Uniform weighting over wider horizons yields no benefit. A small horizon (iโค2, uniform) gives moderate gains and converges faster.

Figure 5: Validation loss for target offset prediction ablations. Exponential weighting over iโค5 is the strongest individual augmentation; uniform weighting over iโค5 provides no benefit.
Compositional Effects and Interference Patterns
Systematic ablation reveals strongly nontrivial effects when composing augmentations Figure 6:
- Noise and Offset: These interfere, with combined models often performing no better than baseline unless the noise rate is extremely low.
- Permutation and Offset: R2L and offset synergize, with the combination matching or slightly improving upon individual minima.
- Noise and Permutation: Random token noise (especially at lower rates) can complement R2L, but masking is detrimental at higher rates.

Figure 6: Systematic 2-cat combinations of the three best individuals. R2L~+~offset synergizes (sky blue, min 3.841). Token noise~+~offset interferes (red-orange, min 3.995). Token noise~+~R2L is intermediate (pink, min 3.887).
Lowering token noise to 5% and applying it with R2L and iโค5 offset prediction yields the best observed minimum (3.805), exceeding any individual or two-way combination. This optimal regularization only occurs within a narrow hyperparameter band Figure 7.

Figure 8: Token noise~ร~R2L combinations across noise rates and types. Random replacement (solid) consistently outperforms masking (also solid, converges earlier). Lower noise rates achieve better minimum loss and require fewer epochs to converge.

Figure 7: All 3-category combinations (solid) vs.\ the best 2-cat base R2L~+~xt+1โ0 exp.\ (dashed). Reducing the noise rate from 15\% to 5\% resolves the interference: Rand.\ 5\%~+~R2L~+~xt+1โ1 exp.\ achieves the overall best minimum of 3.805 at epoch~68.
Decay Phase and Convergence
Learning rate decay preserves the relative ranking of all configurations and generally lowers minima. The 3-category optimal combination (Rand. 5% + R2L + xt+1โ2 exp.) reaches a final minimum of 3.792, the best across the entire experimental suite Figure 9.

Figure 9: Validation loss trajectories for all eight configurations (epoch axis truncated at 150). The 3-category combination (purple, Rand.\ 5\%+R2L+xt+1โ3 exp.) decays from epoch~68 and achieves the lowest decay minimum of 3.792.
Downstream Generalization and Practical Implications
All augmentation strategies that improve validation loss also produce gains in zero-shot downstream evaluation, though absolute improvement margins remain modest at this model scale. There is not a perfect ordinal relationship between validation loss and downstream accuracy, as expected with noisy small models, but enhanced regularization enables the model to maintain generalization in scenarios where standard AR pretraining fails.
From a practical perspective, the results establish that appropriately tuned data augmentations can substantially mitigate the inefficiency of multi-epoch AR pretraining in data-constrained environments. The optimal configuration achieves both lower and later minima than baseline, demonstrating a viable recipe for extending model utility when dataset expansion is infeasible.
Theoretical and Methodological Implications
These findings suggest that the principle of training-time instance-level augmentation, historically central to computer vision, is also effective for regularizing LLMs under repeated exposure to limited data. Notably, certain augmentations (e.g., FIM, masking at high rates) can degrade generalization or interact destructively with others (e.g., token corruption with offset prediction), underlining that regularization by augmentation is highly sensitive to both augmentation composition and parameterization.
Moreover, these strategies leave the AR objective and architecture intact, preserving compatibility with existing LLM infrastructure and making them attractive for deployment within current pretraining pipelines. The study opens multiple research directions: optimal dynamic scheduling of augmentations during training, finer-grained hyperparameter search, investigating scaling laws for augmentation efficacy, and exploring compositional effects in even larger models and diverse corpora.
Conclusion
Training-time data augmentationsโspecifically, token-level random replacement, right-to-left permutation, and target offset predictionโprovide a robust, modular mechanism for regularizing AR LM pretraining when confronted with data scarcity. When combined judiciously, these augmentations recover significant losses due to overfitting, achieving stable and improved validation loss trajectories even with heavy multi-epoch training. The demonstrated improvements, combined with strong downstream transfer, indicate that systematic augmentation should become a first-class technique in data-constrained LM pretraining. Future work will clarify the extent to which these results generalize across model and data regimes, and whether new augmentation categories or adaptive scheduling can further enhance data efficiency in large-scale language modeling.