Generator-Level Delta in Preference Optimization
- Generator-level delta is defined as the capability gap between models that generate the chosen and rejected responses in preference pairs.
- Experimental results demonstrate that increasing generator-level delta enriches reasoning supervision and significantly improves out-of-domain performance.
- This mechanism functions as a dataset-construction knob distinct from sample-level delta, emphasizing model pairing over individual trace quality.
Searching arXiv for the primary paper and closely related uses of “generator-level delta” to ground the article. Generator-level delta is a property of preference-pair construction in preference optimization for LLMs. In the formulation studied by "Decomposing the Delta: What Do Models Actually Learn from Preference Pairs?" (Lee et al., 9 Apr 2026), a preference dataset consists of triples , where is a prompt or problem, is the chosen response, and is the rejected response. Generator-level delta is the capability difference between the model that produces and the model that produces . It is therefore a dataset-construction variable rather than a within-example quality measure. The same work distinguishes it from sample-level delta, which is the judged quality gap within a particular pair of traces. The paper argues that these two deltas play different roles: generator-level delta primarily affects out-of-domain generalization, whereas sample-level delta primarily affects data efficiency (Lee et al., 9 Apr 2026).
1. Definition and conceptual scope
Generator-level delta is defined as “the capability difference between the model that produces chosen responses and the model that produces rejected responses” (Lee et al., 9 Apr 2026). In practical terms, it is instantiated by pairing outputs from a stronger model with outputs from a weaker model. The construction is agnostic to any single symbolic formula for delta itself; instead, delta is interpreted as a property of the generator pair associated with the standard preference tuple
This concept is specific to how preference data are assembled. It does not describe the final answer correctness of a pair, nor does it directly quantify judged trace quality. Those latter properties belong to sample-level analysis. The distinction is central in (Lee et al., 9 Apr 2026): generator-level delta answers who produced the pair, while sample-level delta answers how different the two responses are within that pair.
A compact summary given by the paper is that generator-level delta is a dataset-construction knob, whereas sample-level delta is a pair-selection knob (Lee et al., 9 Apr 2026). This suggests that the former operates at corpus design time, and the latter at corpus filtering or subsampling time.
2. Role in preference optimization
The experimental framing in (Lee et al., 9 Apr 2026) is standard DPO/KTO preference optimization. The paper explicitly trains with both methods. The reported hyperparameters are:
| Method | Learning rate | Key parameters |
|---|---|---|
| DPO | ||
| KTO | 0, desirable/undesirable weights both 1 |
Within this setup, generator-level delta is not a modification to the DPO or KTO loss itself. Instead, it changes the supervision distribution supplied to those objectives. The paper’s interpretation is that larger generator-level delta increases the density and richness of the preference signal (Lee et al., 9 Apr 2026).
The experimental setup for varying generator-level delta fixes the base policy at Nemotron-8B and varies the response-synthesis families across Nemotron and s1 models. The chosen generator may be s1-7B, s1-14B, s1-32B, Nemotron-4B, Nemotron-8B, or Nemotron-49B, while the rejected generator is fixed as s1-3B (Lee et al., 9 Apr 2026). One example given explicitly is a pair with 2 from s1-7B and 3 from s1-3B. Stronger-gap conditions use progressively stronger chosen generators against the same rejected generator (Lee et al., 9 Apr 2026).
The training data source is OpenR1-Math-220k, filtered to 16.5k problems that contain at least one verified-correct and one verified-incorrect trace. Verification is performed by Math-Verify for most samples and by Llama-3.3-70B-Instruct for the rest. Inference uses temperature 0.7 and top-4, with pass@1 for most benchmarks and avg@16 for AMC/AIME tasks (Lee et al., 9 Apr 2026).
3. Experimental instantiation and measured trace quality
The paper studies generator-level delta by varying generator scale and model family while holding the rejected side fixed (Lee et al., 9 Apr 2026). This design makes the capability gap an explicit experimental variable.
The authors then measure how this generator-level design choice affects within-pair quality gaps along five reasoning-quality dimensions:
- factuality
- strategy coherence
- step coherence
- computational precision
- signal-to-noise
These dimensions are scored from 1 to 5 using GPT-OSS-120b as judge (Lee et al., 9 Apr 2026). The reported result is that increasing generator-level delta steadily widens the average sample-level gaps across all five dimensions. For example, over the progression from s1-7B vs s1-3B to Nemotron-49B vs s1-3B, step coherence rises from 0.96 to 2.08, factuality from 1.06 to 1.77, and strategy coherence from 0.92 to 1.81 (Lee et al., 9 Apr 2026).
This is important because it links the dataset-level generator choice to the pair-level informational contrast. The paper presents the widening of these judged quality gaps as the mechanism behind improved downstream generalization (Lee et al., 9 Apr 2026). A plausible implication is that generator-level delta is useful not merely because stronger models are more often correct, but because they alter multiple dimensions of reasoning quality simultaneously.
4. Effects on downstream performance
The downstream findings are explicitly asymmetric across benchmark families. In-domain math performance mostly saturates as generator-level delta increases, whereas out-of-domain performance improves steadily with larger generator-level delta, especially on STEM and code tasks (Lee et al., 9 Apr 2026).
The benchmark suite includes:
| Category | Benchmarks |
|---|---|
| In-domain math | MATH-500, GSM8K, AMC23 |
| Harder math | Minerva-Math, OlympiadBench, AIME24, AIME25 |
| Out-of-domain | MMLU-Pro, TheoremQA, LiveCodeBench |
The paper states that larger capability gaps “consistently translate into stronger out-of-domain generalization” (Lee et al., 9 Apr 2026). The strongest gains are reported on LiveCodeBench, with more modest but still positive effects on MMLU-Pro and TheoremQA (Lee et al., 9 Apr 2026). By contrast, the saturation on in-domain math is attributed to the fact that the base model is already fairly strong on math, making additional improvement difficult to extract (Lee et al., 9 Apr 2026).
This yields a specific empirical interpretation of generator-level delta: its main value is transfer rather than raw specialization. The evidence in (Lee et al., 9 Apr 2026) suggests that larger capability gaps do not simply overfit the training distribution more strongly; instead, they improve general reasoning performance outside the original domain, particularly for STEM and code.
5. Relation to sample-level delta and correctness
A central result of (Lee et al., 9 Apr 2026) is that generator-level delta should not be conflated with sample-level delta. Generator-level delta concerns the identity and capability gap of the models that generated 5 and 6. Sample-level delta concerns the judged difference within an individual pair. The two are related but not interchangeable: increasing generator-level delta tends to widen sample-level quality gaps, yet sample-level delta can also be exploited independently as a filtering signal (Lee et al., 9 Apr 2026).
The paper reports that selecting pairs with larger sample-level differences can match or outperform the full 16.5k set using far fewer examples. Using only the top 5k highest-delta pairs can match or exceed the full set, and even 1k pairs gives consistent gains. On hard math, the full set gets about 60%, while top-5k subsets are within about 7 of that. Among the quality criteria, step coherence is the most consistently useful selection signal (Lee et al., 9 Apr 2026).
The authors also test whether the delta signal is primarily final-answer correctness. It is not. Using Nemotron-4B-generated traces and DPO on Nemotron-8B, they train on all four correctness combinations—correct→incorrect, correct→correct, incorrect→incorrect, and incorrect→correct. All four combinations improve over the base model on in-domain math, with scores ranging roughly 71.9% to 73.0% versus a baseline of about 71.1% (Lee et al., 9 Apr 2026). Even incorrect→correct and incorrect→incorrect pairs help.
This result is conceptually important for generator-level delta. Since helpful signal persists even when both sides are incorrect, the effective supervision cannot be reduced to outcome labels alone. The paper explicitly argues that the preference signal is more about “how the reasoning is done” than just “whether the final answer is correct” (Lee et al., 9 Apr 2026). This suggests that generator-level delta matters because it changes the process-level contrast available to the learner.
6. Interpretation, mechanism, and practical recipe
The intuition given in (Lee et al., 9 Apr 2026) is that a stronger chosen generator does not merely produce more correct outputs. It also tends to produce traces with better logical structure, fewer computation errors, and more relevant reasoning steps. Larger generator-level delta therefore increases the richness of the contrast between chosen and rejected traces (Lee et al., 9 Apr 2026).
This mechanism is consistent with the measured widening of factuality, strategy coherence, step coherence, computational precision, and signal-to-noise gaps (Lee et al., 9 Apr 2026). The paper treats these widening gaps as the operational content of improved supervision quality. In that sense, generator-level delta is not a latent abstraction detached from data; it manifests through systematic changes in the structure of the trace pairs.
The practical recommendation in (Lee et al., 9 Apr 2026) is a two-part recipe:
- maximize generator-level delta when constructing preference pairs, by pairing outputs from models with a large capability gap;
- exploit sample-level delta to select the most informative pairs, especially those with large step coherence differences.
This recommendation is deliberately split because the two deltas solve different problems. Larger generator-level delta mainly improves out-of-domain generalization. Larger sample-level delta mainly improves data efficiency (Lee et al., 9 Apr 2026).
A common misconception would be to treat generator-level delta as identical to correctness delta or to assume it is valuable only when both generators come from the same family. The paper does not support either view. It varies both scale and model family, and its correctness-combination experiments show that the useful signal is not exhausted by answer correctness (Lee et al., 9 Apr 2026).
7. Broader usage of the term “generator-level delta”
Outside preference optimization, “generator-level delta” appears in other technical literatures with different meanings. In high-energy physics, for example, "Reweighting and Analysing Event Generator Systematics by Neural Networks on High-Level Features" (Furuichi et al., 3 Mar 2025) uses the term to denote the discrepancy between jet distributions produced by different Monte Carlo event generators, specifically Pythia and Herwig. There, the objective is density-ratio estimation and reweighting between source and target generator distributions, not preference-pair construction (Furuichi et al., 3 Mar 2025).
In software testing, "Validity-Preserving Delta Debugging via Generator Trace Reduction" (Ren et al., 2024) uses a generator-level delta idea to mean reducing the execution trace of the input generator rather than reducing the generated artifact directly. The motivation is validity preservation under semantic constraints, again unrelated to preference optimization (Ren et al., 2024).
These usages share a structural theme: the relevant delta is defined at the level of the generative mechanism rather than only at the level of final outputs. In the preference-learning setting of (Lee et al., 9 Apr 2026), this means the capability gap between the models that generate 8 and 9. In other domains, it may mean generator-induced distribution shift or modifications to generator executions. This suggests that “generator-level delta” is best understood as a meta-level contrast residing in the producing system, with task-specific semantics determined by the surrounding framework.
In the specific sense established by (Lee et al., 9 Apr 2026), however, generator-level delta is a design variable for preference data construction. Its significance lies in the empirical claim that larger generator capability gaps enrich reasoning supervision and thereby strengthen out-of-domain generalization, while remaining distinct from pairwise filtering criteria such as sample-level delta (Lee et al., 9 Apr 2026).