RL2ML: Finite-Rollout Surrogate Objectives
- RL2ML is a family of finite‐rollout surrogate objectives designed for correctness‐based reinforcement learning with verifiable rewards, decoupling the expected gradient from finite sample updates.
- It formulates a continuous one-parameter family indexed by γ, interpolating between ordinary RL (γ=0), maximum likelihood (γ=1), and beyond, with explicit unbiased gradient estimators.
- The framework clarifies how finite rollout budgets influence group-level update scales and estimator variance, providing a nuanced perspective on objective design in RLVR.
RL2ML denotes a family of finite-rollout surrogate objectives for correctness-based reinforcement learning with verifiable rewards (RLVR), introduced to separate the population-level objective optimized in expectation from the stochastic update geometry induced by finite rollout groups (Zheng, 28 May 2026). In this formulation, language-model training from binary correctness signals is not treated as a single fixed point between ordinary reinforcement learning and maximum likelihood, but as a continuous one-parameter family indexed by , with exact finite-rollout estimators aligned to the corresponding surrogate objective (Zheng, 28 May 2026). The framework is motivated by the observation that in RLVR, the objective optimized in expectation and the updates produced under a fixed rollout budget are often conflated, even though finite-group success counts can induce qualitatively different optimization behavior (Zheng, 28 May 2026).
1. RL2ML problem formulation
RL2ML is defined in the setting of correctness-based RL with verifiable rewards. For each prompt , a model samples a latent rollout , applies deterministic decoding , and receives binary reward
For a fixed prompt , the model success probability is
The score function is
If rollouts are sampled independently, the paper defines
0
and, when 1,
2
It further introduces
3
with the identities
4
The central problem is that RLVR training uses finite rollout groups, so optimization is governed not only by the expected gradient but also by how observed groups with different empirical success count 5 are scaled after sampling (Zheng, 28 May 2026). RL2ML treats this distinction as fundamental rather than incidental. This suggests that finite-budget training cannot be fully characterized by population-level notation such as 6 or 7 alone.
2. Objective family from reinforcement learning to maximum likelihood
RL2ML defines a one-parameter family of prompt-level surrogate objectives: 8 The antiderivative is
9
Hence the untruncated population gradient is
0
This produces a continuous interpolation across three regimes (Zheng, 28 May 2026):
- 1: ordinary RL, since
2
- 3: maximum likelihood, since
4
- 5: “beyond-maximum-likelihood” objectives that weight low-success prompts more strongly than 6 (Zheng, 28 May 2026).
For finite rollout budget 7, the paper defines a rollout-aligned truncated surrogate: 8 with population-level weight
9
and gradient
0
This finite-rollout objective is the relevant one for training under a fixed rollout budget, rather than the asymptotic 1 expression alone (Zheng, 28 May 2026).
3. Exactly unbiased finite-rollout estimator
A central contribution of RL2ML is a closed-form, exactly unbiased gradient estimator for the finite-rollout truncated surrogate (Zheng, 28 May 2026). The paper first considers estimators of the form
2
defines
3
and uses the Bernstein basis
4
It then shows
5
This Bernstein representation is the mechanism by which a 6-only estimator induces a specific population weight on 7 (Zheng, 28 May 2026).
For RL2ML, the Bernstein coefficients are
8
The corresponding group-level update scale is
9
The estimator is then
0
and satisfies
1
The paper emphasizes that this is exact unbiasedness for the finite-rollout surrogate objective itself, not merely for an asymptotic population target (Zheng, 28 May 2026).
RL2ML also gives a control-variate form: 2 with sequence-level advantage
3
At 4,
5
which recovers MaxRL’s control-variate form (Zheng, 28 May 2026).
4. Group-level update geometry and the 6 transition
RL2ML distinguishes population weighting from group-level update geometry by focusing on the empirical success count 7 of a rollout group (Zheng, 28 May 2026). The stochastic update can be written as
8
so 9 directly controls the expected update magnitude and within-group variance conditional on the observed group.
For RL2ML, the conditional moments are
0
and
1
Thus 2 is the operative quantity that determines how observed groups are amplified or attenuated (Zheng, 28 May 2026).
The paper derives
3
which reveals a subcritical-supercritical update-scale transition at
4
This yields four regimes (Zheng, 28 May 2026):
- RL endpoint (5):
6
- Subcritical regime (7): 8 is increasing in 9, and
0
- ML boundary / MaxRL boundary (1):
2
- Supercritical regime (3): 4 is decreasing in 5, and
6
This distinction is one of RL2ML’s main conceptual claims: 7 is a structural boundary in the sample-level geometry of finite-rollout optimization, not merely the point corresponding to 8 in population notation (Zheng, 28 May 2026).
5. Calibrated metric gain and exact variance decomposition
RL2ML argues that the best surrogate objective is determined neither by proximity to maximum likelihood nor by the population-level weight alone (Zheng, 28 May 2026). Instead, it depends jointly on the evaluation metric, local prompt sensitivity, and estimator variance.
For a prompt-separable validation metric
9
the paper defines
0
1
After calibrating all candidate objectives to the same update length 2,
3
the first-order metric-gain criterion is
4
The local gain is
5
This formalizes the claim that the preferred surrogate under a fixed rollout budget must be evaluated relative to a concrete downstream metric, not only relative to ML-like weighting (Zheng, 28 May 2026).
The variance analysis is equally explicit. With
6
the paper proves
7
The corresponding conditional mean-squared deviation is
8
The two terms separate count variance from within-success variance (Zheng, 28 May 2026). A plausible implication is that supercritical objectives can improve hard-prompt emphasis while simultaneously incurring a distinct finite-rollout noise cost.
The remaining degree of freedom is therefore cast as a one-dimensional optimization problem: 9 where
0
This is presented as a principled alternative to treating objective choice as an unconstrained hyperparameter search (Zheng, 28 May 2026).
6. Position within the broader RL-for-ML landscape
RL2ML is primarily a theoretical and analytical contribution rather than a benchmark-heavy empirical paper (Zheng, 28 May 2026). Its experiments are conceptual figures, stylized finite-horizon selection plots, and implementation guidance for Verl-style RLVR systems, rather than large real-task comparisons (Zheng, 28 May 2026). The paper’s main value is to reinterpret objective design in finite-rollout RLVR through estimator-objective alignment, group-level success-count geometry, and variance-aware metric optimization (Zheng, 28 May 2026).
In a broader methodological sense, RL2ML belongs to a line of work in which reinforcement learning is used to shape ML systems rather than only downstream task outputs. That broader pattern appears in several distinct settings:
| Work | RL target | Core role of RL |
|---|---|---|
| RL2ML | finite-rollout surrogate objective design | aligns estimator and objective under fixed rollout budget |
| ReMix | discrete LoRA routing | trains a non-differentiable router with REINFORCE/RLOO (Qiu et al., 10 Mar 2026) |
| R2-Reasoner | task decomposition and model allocation | optimizes a multi-model routing policy with GRPO (Shao et al., 6 Jun 2025) |
| Retrv-R1 | retrieval reasoning and inspection actions | trains correctness-efficiency retrieval behavior with GRPO (Zhu et al., 3 Oct 2025) |
These neighboring examples differ in mechanism, but they share a structural feature with RL2ML: reinforcement learning is used to optimize architectural, routing, or systems-level decisions rather than only a final response distribution. This suggests that RL2ML’s emphasis on finite-sample update geometry is relevant beyond correctness-based RLVR, especially in settings where rollout groups, sparse rewards, and budgeted stochastic updates are central.
RL2ML’s main technical takeaway is that finite-rollout objective design should not be reduced to a binary choice between RL and maximum likelihood. Under a fixed rollout budget, the effective training signal depends on the surrogate family
1
the observed success count 2, the induced group-level scale 3, the target evaluation metric, and the resulting estimator variance (Zheng, 28 May 2026). Within that framework, 4 remains important as the maximum-likelihood boundary and the point where 5 becomes constant across successful groups, but it is not presented as a universal optimum (Zheng, 28 May 2026).