---
title: Generator-Level Delta in Preference Optimization
url: https://www.emergentmind.com/topics/generator-level-delta
type: topic
---

# Generator-Level Delta in Preference Optimization

Searching arXiv for the primary paper and closely related uses of “generator-level delta” to ground the article.
Generator-level delta is a property of preference-pair construction in preference optimization for language models. In the formulation studied by "Decomposing the Delta: What Do Models Actually Learn from Preference Pairs?" [2604.08723], a preference dataset consists of triples \((x, y^{+}, y^{-})\), where \(x\) is a prompt or problem, \(y^{+}\) is the chosen response, and \(y^{-}\) is the rejected response. Generator-level delta is the capability difference between the model that produces \(y^{+}\) and the model that produces \(y^{-}\). It is therefore a dataset-construction variable rather than a within-example quality measure. The same work distinguishes it from sample-level delta, which is the judged quality gap within a particular pair of traces. The paper argues that these two deltas play different roles: generator-level delta primarily affects out-of-domain generalization, whereas sample-level delta primarily affects data efficiency [2604.08723].

## 1. Definition and conceptual scope

Generator-level delta is defined as “the capability difference between the model that produces chosen responses and the model that produces rejected responses” [2604.08723]. In practical terms, it is instantiated by pairing outputs from a stronger model with outputs from a weaker model. The construction is agnostic to any single symbolic formula for delta itself; instead, delta is interpreted as a property of the generator pair associated with the standard preference tuple
\[
(x, y^{+}, y^{-}).
\]

This concept is specific to how preference data are assembled. It does not describe the final answer correctness of a pair, nor does it directly quantify judged trace quality. Those latter properties belong to sample-level analysis. The distinction is central in [2604.08723]: generator-level delta answers who produced the pair, while sample-level delta answers how different the two responses are within that pair.

A compact summary given by the paper is that generator-level delta is a dataset-construction knob, whereas sample-level delta is a pair-selection knob [2604.08723]. This suggests that the former operates at corpus design time, and the latter at corpus filtering or subsampling time.

## 2. Role in preference optimization

The experimental framing in [2604.08723] is standard DPO/KTO preference optimization. The paper explicitly trains with both methods. The reported hyperparameters are:

| Method | Learning rate | Key parameters |
|---|---:|---|
| DPO | \(5 \times 10^{-8}\) | \(\beta = 0.2\) |
| KTO | \(5 \times 10^{-8}\) | \(\beta = 0.5\), desirable/undesirable weights both \(1.0\) |

Within this setup, generator-level delta is not a modification to the DPO or KTO loss itself. Instead, it changes the supervision distribution supplied to those objectives. The paper’s interpretation is that larger generator-level delta increases the density and richness of the preference signal [2604.08723].

The experimental setup for varying generator-level delta fixes the base policy at Nemotron-8B and varies the response-synthesis families across Nemotron and s1 models. The chosen generator may be s1-7B, s1-14B, s1-32B, Nemotron-4B, Nemotron-8B, or Nemotron-49B, while the rejected generator is fixed as s1-3B [2604.08723]. One example given explicitly is a pair with \(y^{+}\) from s1-7B and \(y^{-}\) from s1-3B. Stronger-gap conditions use progressively stronger chosen generators against the same rejected generator [2604.08723].

The training data source is OpenR1-Math-220k, filtered to 16.5k problems that contain at least one verified-correct and one verified-incorrect trace. Verification is performed by Math-Verify for most samples and by Llama-3.3-70B-Instruct for the rest. Inference uses temperature 0.7 and top-\(p=1.0\), with pass@1 for most benchmarks and avg@16 for AMC/AIME tasks [2604.08723].

## 3. Experimental instantiation and measured trace quality

The paper studies generator-level delta by varying generator scale and model family while holding the rejected side fixed [2604.08723]. This design makes the capability gap an explicit experimental variable.

The authors then measure how this generator-level design choice affects within-pair quality gaps along five reasoning-quality dimensions:

1. factuality  
2. strategy coherence  
3. step coherence  
4. computational precision  
5. signal-to-noise  

These dimensions are scored from 1 to 5 using GPT-OSS-120b as judge [2604.08723]. The reported result is that increasing generator-level delta steadily widens the average sample-level gaps across all five dimensions. For example, over the progression from s1-7B vs s1-3B to Nemotron-49B vs s1-3B, step coherence rises from 0.96 to 2.08, factuality from 1.06 to 1.77, and strategy coherence from 0.92 to 1.81 [2604.08723].

This is important because it links the dataset-level generator choice to the pair-level informational contrast. The paper presents the widening of these judged quality gaps as the mechanism behind improved downstream generalization [2604.08723]. A plausible implication is that generator-level delta is useful not merely because stronger models are more often correct, but because they alter multiple dimensions of reasoning quality simultaneously.

## 4. Effects on downstream performance

The downstream findings are explicitly asymmetric across benchmark families. In-domain math performance mostly saturates as generator-level delta increases, whereas out-of-domain performance improves steadily with larger generator-level delta, especially on STEM and code tasks [2604.08723].

The benchmark suite includes:

| Category | Benchmarks |
|---|---|
| In-domain math | MATH-500, GSM8K, AMC23 |
| Harder math | Minerva-Math, OlympiadBench, AIME24, AIME25 |
| Out-of-domain | MMLU-Pro, TheoremQA, LiveCodeBench |

The paper states that larger capability gaps “consistently translate into stronger out-of-domain generalization” [2604.08723]. The strongest gains are reported on LiveCodeBench, with more modest but still positive effects on MMLU-Pro and TheoremQA [2604.08723]. By contrast, the saturation on in-domain math is attributed to the fact that the base model is already fairly strong on math, making additional improvement difficult to extract [2604.08723].

This yields a specific empirical interpretation of generator-level delta: its main value is transfer rather than raw specialization. The evidence in [2604.08723] suggests that larger capability gaps do not simply overfit the training distribution more strongly; instead, they improve general reasoning performance outside the original domain, particularly for STEM and code.

## 5. Relation to sample-level delta and correctness

A central result of [2604.08723] is that generator-level delta should not be conflated with sample-level delta. Generator-level delta concerns the identity and capability gap of the models that generated \(y^{+}\) and \(y^{-}\). Sample-level delta concerns the judged difference within an individual pair. The two are related but not interchangeable: increasing generator-level delta tends to widen sample-level quality gaps, yet sample-level delta can also be exploited independently as a filtering signal [2604.08723].

The paper reports that selecting pairs with larger sample-level differences can match or outperform the full 16.5k set using far fewer examples. Using only the top 5k highest-delta pairs can match or exceed the full set, and even 1k pairs gives consistent gains. On hard math, the full set gets about 60%, while top-5k subsets are within about \(\pm 0.65\%\) of that. Among the quality criteria, step coherence is the most consistently useful selection signal [2604.08723].

The authors also test whether the delta signal is primarily final-answer correctness. It is not. Using Nemotron-4B-generated traces and DPO on Nemotron-8B, they train on all four correctness combinations—correct→incorrect, correct→correct, incorrect→incorrect, and incorrect→correct. All four combinations improve over the base model on in-domain math, with scores ranging roughly 71.9% to 73.0% versus a baseline of about 71.1% [2604.08723]. Even incorrect→correct and incorrect→incorrect pairs help.

This result is conceptually important for generator-level delta. Since helpful signal persists even when both sides are incorrect, the effective supervision cannot be reduced to outcome labels alone. The paper explicitly argues that the preference signal is more about “how the reasoning is done” than just “whether the final answer is correct” [2604.08723]. This suggests that generator-level delta matters because it changes the process-level contrast available to the learner.

## 6. Interpretation, mechanism, and practical recipe

The intuition given in [2604.08723] is that a stronger chosen generator does not merely produce more correct outputs. It also tends to produce traces with better logical structure, fewer computation errors, and more relevant reasoning steps. Larger generator-level delta therefore increases the richness of the contrast between chosen and rejected traces [2604.08723].

This mechanism is consistent with the measured widening of factuality, strategy coherence, step coherence, computational precision, and signal-to-noise gaps [2604.08723]. The paper treats these widening gaps as the operational content of improved supervision quality. In that sense, generator-level delta is not a latent abstraction detached from data; it manifests through systematic changes in the structure of the trace pairs.

The practical recommendation in [2604.08723] is a two-part recipe:

1. maximize generator-level delta when constructing preference pairs, by pairing outputs from models with a large capability gap;  
2. exploit sample-level delta to select the most informative pairs, especially those with large step coherence differences.  

This recommendation is deliberately split because the two deltas solve different problems. Larger generator-level delta mainly improves out-of-domain generalization. Larger sample-level delta mainly improves data efficiency [2604.08723].

A common misconception would be to treat generator-level delta as identical to correctness delta or to assume it is valuable only when both generators come from the same family. The paper does not support either view. It varies both scale and model family, and its correctness-combination experiments show that the useful signal is not exhausted by answer correctness [2604.08723].

## 7. Broader usage of the term “generator-level delta”

Outside preference optimization, “generator-level delta” appears in other technical literatures with different meanings. In high-energy physics, for example, "Reweighting and Analysing Event Generator Systematics by Neural Networks on High-Level Features" [2503.01452] uses the term to denote the discrepancy between jet distributions produced by different Monte Carlo event generators, specifically Pythia and Herwig. There, the objective is density-ratio estimation and reweighting between source and target generator distributions, not preference-pair construction [2503.01452].

In software testing, "Validity-Preserving Delta Debugging via Generator Trace Reduction" [2402.04623] uses a generator-level delta idea to mean reducing the execution trace of the input generator rather than reducing the generated artifact directly. The motivation is validity preservation under semantic constraints, again unrelated to preference optimization [2402.04623].

These usages share a structural theme: the relevant delta is defined at the level of the generative mechanism rather than only at the level of final outputs. In the preference-learning setting of [2604.08723], this means the capability gap between the models that generate \(y^{+}\) and \(y^{-}\). In other domains, it may mean generator-induced distribution shift or modifications to generator executions. This suggests that “generator-level delta” is best understood as a meta-level contrast residing in the producing system, with task-specific semantics determined by the surrounding framework.

In the specific sense established by [2604.08723], however, generator-level delta is a design variable for preference data construction. Its significance lies in the empirical claim that larger generator capability gaps enrich reasoning supervision and thereby strengthen out-of-domain generalization, while remaining distinct from pairwise filtering criteria such as sample-level delta [2604.08723].

Source: https://www.emergentmind.com/topics/generator-level-delta