Papers
Topics
Authors
Recent
Search
2000 character limit reached

Not All Synthetic Data Is Yours to Learn From

Published 29 May 2026 in cs.CL, cs.AI, and cs.LG | (2605.31126v1)

Abstract: Can a LLM improve from plain text sampled from itself, with no prompts, no teacher, no verifier, and no reward model? Yes, but only when the synthetic corpus is compatible with the student, a relational property of the source-student pair rather than an intrinsic property of the data. We call this the latent capability resurfacing hypothesis: weak self-training can amplify capabilities already present in the pretrained model, but only under this compatibility condition. We study this in the minimal setting of prompt-free unconditional self-training, where base LLMs are fine-tuned on text generated from the BOS token alone, with no task specification or external supervision. We report three findings. First, synthetic utility is relational rather than intrinsic: self-generated data is the most effective source, same-lineage transfer outperforms stronger but differently trained sources, and cross-family transfer is substantially weaker. Second, common intrinsic proxies fail: neither benchmark-level semantic similarity nor average per-token likelihood under the student predicts which corpora help. Third, this regime produces a surprising byproduct. In controlled Pythia experiments, capability and verbatim memorization decouple: benchmark utility is preserved or improved while held-out exact-match extraction drops by over 95 percent, with no forget set, privacy objective, or targeted unlearning. Together, these results suggest that prompt-free self-training works by amplifying what the student already knows, not by importing structure from the data. They also reveal a regime in which capability and verbatim memorization can be separated without any explicit unlearning objective.

Summary

  • The paper demonstrates that BOS-only self-training selectively improves reasoning, math, and code performance in Qwen2.5-0.5B, while gains are transient or absent in LLaMA-3.2-1B and are not reproduced by a real-text replay baseline.
  • The paper finds that synthetic-data utility is relational: model lineage and student–source compatibility predict transfer better than source capability, benchmark similarity, or mean student likelihood, with Qwen self-data outperforming stronger or cross-family alternatives.
  • The paper shows that favorable self-training can preserve or improve benchmark accuracy while reducing held-out memorization extraction by roughly 94–97% for Pythia text sequences and about 48–50% for code sequences, without an explicit unlearning objective.

The setting and the hypothesis

This paper studies the most stripped-down version of self-training that can be defined: a base LLM is fine-tuned on plain text sampled unconditionally from another model's BOS token, with no prompts, no task specification, no verifier, no reward model, and no teacher. The authors use this minimal regime to test what they call the latent capability resurfacing hypothesis, which makes three linked claims: (i) pretrained models contain useful capabilities not fully expressed by their base behavior; (ii) weak synthetic fine-tuning can surface these capabilities, but only when the corpus is compatible with the student — a relational property of the source–student pair rather than an intrinsic property of the data; and (iii) in favorable regimes the update amplifies distributed task structure rather than sequence-specific recall.

The motivation comes from an apparent tension in the literature. Recursive self-consumption is associated with distributional drift and model collapse (Alemohammad et al., 2023), and post-training pipelines typically treat unverified synthetic data as unreliable (Feng et al., 2024). Yet several recent results show weak, unstructured signals helping anyway: unfiltered self-distillation improves code generation (Zhang et al., 1 Apr 2026), reward-free self-training helps reasoning (Li et al., 21 Oct 2025), and training-free iterative sampling nearly matches RL-post-trained models (Karan et al., 16 Oct 2025). Each of these settings introduces confounds through prompts, rewards, or inference-time orchestration; the BOS-only design removes all of them by construction, so any gain must come from the interaction between the pretrained student and a weak self-generated signal.

Transient, model-dependent gains from unconditional self-generated text

The core protocol generates roughly 5M-token corpora at temperatures τ∈{0.75,1.0,1.25}\tau \in \{0.75, 1.0, 1.25\} with no truncation, applies 8-gram decontamination against all benchmark test splits, draws three independent stratified subsets per corpus to estimate subset sensitivity, and fine-tunes base models for 40 epochs at learning rate 10−610^{-6} with full-parameter AdamW. A matched Common Corpus replay baseline controls for generic low-learning-rate regularization.

On Qwen2.5-0.5B, self-generated data produces clear gains on structured reasoning, math, and code: sustained positive deltas on ARC-Challenge and HellaSwag saturating within 20–30 epochs, a pronounced early GSM8K spike under τ=1.25\tau=1.25 that decays toward baseline (the transient-then-degrade signature of a weak-signal regime), and modest HumanEval improvements across temperatures. The replay baseline fails to reproduce these gains — flat or negative on GSM8K, flat on Minerva-MATH, degrading on HumanEval — while matching synthetic data only on MMLU and TruthfulQA. The implication is that the effect is specific to structured-reasoning and code capabilities concentrated in the model's own high-probability modes, not a generic regularization artifact.

The identical protocol on LLaMA-3.2-1B does not reproduce this pattern. GSM8K and Minerva-MATH show no improvement under any condition, HumanEval degrades substantially, and only comprehension-style benchmarks improve modestly — often with the replay baseline matching or exceeding synthetic data. This qualitative task-type split implies that the set of capabilities amenable to resurfacing is fixed by the student's pretraining, mirroring the Qwen-versus-LLaMA asymmetry reported for spurious-reward RLVR (Shao et al., 12 Jun 2025).

The paper also rules out benchmark contamination as an explanation. Embedding every synthetic sample and benchmark item with llama-embed-nemotron-8b and partitioning the Qwen τ=1.25\tau=1.25 corpus by maximum cosine similarity to GSM8K items (below vs. above 0.35) yields nearly indistinguishable GSM8K learning curves: samples with no measurable semantic relationship to GSM8K produce the same improvement as nominally higher-similarity samples. Benchmark proximity therefore does not explain the gains.

Synthetic utility is relational, not intrinsic

Fixing the student to Qwen2.5-0.5B and varying the source produces a clear compatibility hierarchy: self-generated data is strongest; Qwen2.5-7B (larger, same lineage) ranks second; Qwen3-8B — larger and more capable but trained with a different recipe — transfers worse than Qwen2.5-7B; and LLaMA-3.2-1B (same scale, different family) is weakest and often harmful. Pretraining lineage predicts transfer better than raw source capability. A symmetric experiment with a LLaMA student trained on Qwen teacher data confirms the interpretation: Qwen data is not a generic upgrade, since it fails to outperform LLaMA-self on the benchmarks where LLaMA gains at all and actively underperforms on HellaSwag and ARC-Challenge.

Four natural intrinsic proxies fail to explain this ranking:

Proxy Result
Source capability Qwen3-8B transfers worse than smaller Qwen2.5-7B
Student-independent corpus quality Symmetric cross-family experiments rule it out
Benchmark-level semantic similarity Low-similarity subsets match high-similarity ones on GSM8K
Mean likelihood under the student Own and cross corpora converge to near-identical NLL yet diverge downstream

The likelihood result is the sharpest. At τ=1.25\tau=1.25, the Qwen-own and LLaMA-cross corpora have mean per-token NLL of 9.10 and 9.01 respectively under the same scorer — essentially identical — yet one substantially improves structured reasoning while the other leaves the student flat or worse. The mirror analysis under the LLaMA scorer shows the same convergence (μ=8.88\mu = 8.88 vs. $8.93$). This directly qualifies mechanistic accounts of RLVR that emphasize bias toward high-probability pretraining priors (Shao et al., 12 Jun 2025): within compatible pairings, higher-NLL samples at τ=1.25\tau=1.25 can outperform lower-NLL samples at τ=0.75\tau=0.75. What matters is whether the corpus induces useful update directions for the particular student. The authors are explicit that they do not directly measure gradient alignment or update geometry, so compatibility remains an empirically supported explanatory hypothesis rather than a demonstrated mechanism.

Self-training improves the utility–extractability frontier

If resurfacing amplifies distributed structure rather than sequence-specific recall, favorable self-training should reduce verbatim extraction. Testing this on Pythia-1B and Pythia-6.9B — whose documented pretraining corpus makes them standard memorization testbeds (Carlini et al., 2022) — using the prefix-extraction attack on 40,000 held-out sequences from Wikipedia, Enron, Pile-CC, and GitHub yields the paper's most striking numbers. At learning rate 10−510^{-5}:

Model / domain Base extraction Post-training Reduction
Pythia-1B, text ~184 sequences ~0 ~97%
Pythia-6.9B, text ~540 sequences ~34 ~94%
Pythia-1B, code ~1900 ~950 ~50%
Pythia-6.9B, code ~2660 ~1390 ~48%

Meanwhile capability is preserved or improved: Pythia-6.9B at 10−610^{-6}0 gains roughly +2.5 on ARC-Challenge and +1.5 on HellaSwag while memorization collapses. Teacher-forced log-probability of true continuations drops under exactly the conditions where extraction drops, confirming a genuine redistribution of probability mass rather than a decoding artifact. Across every condition where either quantity shifts appreciably, capability and memorization move in opposite directions; the two sub-threshold conditions bound the effect (both quantities barely move at Pythia-6.9B/10−610^{-6}1, while Pythia-1B/10−610^{-6}2 shows the transient-then-degrade pattern with memorization still falling steadily).

This co-occurrence rules out rehearsal of memorized data as the mechanism behind capability gains — if rehearsal were responsible, extraction would rise alongside capability. It also contrasts sharply with targeted unlearning methods (gradient ascent [2202.xxxx-style approaches such as jang2022knowledge], preference optimization on forget sets (Maini et al., 2024)), which typically degrade utility, and with collapse-triggered unlearning that preserves utility at best (Scholten et al., 6 Jul 2025). Here no forget set is specified, no privacy objective is optimized, and no utility is sacrificed; the decoupling is emergent. The authors claim this is the first regime in which plain untargeted self-training jointly improves capability and reduces memorization without any specified forget set or privacy objective — a strong claim, though its scope is confined to the controlled Pythia setting described below.

Limitations and open questions

Several limitations bear directly on how far these results generalize. First, the relational-utility hierarchy rests on a small set of open-weight families and scales (Qwen2.5/3, LLaMA-3.2, Pythia), so the compatibility ordering should not be read as universal across all pretrained models. Second, compatibility itself is inferred from outcomes; gradient alignment and the geometry of induced updates are not measured, leaving the mechanism formally undemonstrated. Third, the memorization analysis uses held-out exact-match extraction and true-continuation likelihood on documented Pythia sequences — stronger than greedy extraction alone, but not a complete privacy audit, and the Pythia models are small relative to production systems where memorization dynamics may differ. Finally, evaluation relies primarily on static benchmarks; whether resurfacing and utility–extractability patterns hold under dynamic, contamination-resistant evaluations remains open. Two further questions the paper leaves unanswered: whether a predictive measure of student–source compatibility can be constructed without running the fine-tuning itself, and whether the memorization reduction persists under longer-horizon recursive training rather than single-generation self-training.

Conclusion

In the strictly minimal BOS-only regime, prompt-free self-training works selectively and relationally: self-generated and close-lineage corpora help, source capability, benchmark proximity, and mean student likelihood do not predict utility, and the recoverable gains are bounded by what the student's pretraining latently supports. The same regime separates capability from verbatim extractability — over 95% reductions in held-out exact-match extraction with preserved or improved benchmark accuracy, confirmed by teacher-forced log-probabilities — without any explicit unlearning objective. The practical implication is that synthetic post-training should be designed around student–source compatibility rather than synthetic data quality alone, and the theoretical implication is that weak self-training amplifies distributed latent structure rather than importing new information from the corpus.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.