Papers
Topics
Authors
Recent
Search
2000 character limit reached

RL2ML: Finite-Rollout Surrogate Objectives

Updated 5 July 2026
  • RL2ML is a family of finite‐rollout surrogate objectives designed for correctness‐based reinforcement learning with verifiable rewards, decoupling the expected gradient from finite sample updates.
  • It formulates a continuous one-parameter family indexed by γ, interpolating between ordinary RL (γ=0), maximum likelihood (γ=1), and beyond, with explicit unbiased gradient estimators.
  • The framework clarifies how finite rollout budgets influence group-level update scales and estimator variance, providing a nuanced perspective on objective design in RLVR.

RL2ML denotes a family of finite-rollout surrogate objectives for correctness-based reinforcement learning with verifiable rewards (RLVR), introduced to separate the population-level objective optimized in expectation from the stochastic update geometry induced by finite rollout groups (Zheng, 28 May 2026). In this formulation, language-model training from binary correctness signals is not treated as a single fixed point between ordinary reinforcement learning and maximum likelihood, but as a continuous one-parameter family indexed by γ\gamma, with exact finite-rollout estimators aligned to the corresponding surrogate objective (Zheng, 28 May 2026). The framework is motivated by the observation that in RLVR, the objective optimized in expectation and the updates produced under a fixed rollout budget NN are often conflated, even though finite-group success counts can induce qualitatively different optimization behavior (Zheng, 28 May 2026).

1. RL2ML problem formulation

RL2ML is defined in the setting of correctness-based RL with verifiable rewards. For each prompt xx, a model samples a latent rollout z∼mθ(⋅∣x)z \sim m_\theta(\cdot \mid x), applies deterministic decoding y=f(z)y=f(z), and receives binary reward

r(x,z):=I{f(z)=y∗(x)}.r(x,z) := I\{f(z)=y^*(x)\}.

For a fixed prompt xx, the model success probability is

pθ(x):=Ez∼mθ(⋅∣x)[r(x,z)].p_\theta(x) := E_{z\sim m_\theta(\cdot\mid x)}[r(x,z)].

The score function is

S(x,z):=∇θlog⁡mθ(z∣x).S(x,z):=\nabla_\theta\log m_\theta(z\mid x).

If NN rollouts are sampled independently, the paper defines

NN0

and, when NN1,

NN2

It further introduces

NN3

with the identities

NN4

(Zheng, 28 May 2026).

The central problem is that RLVR training uses finite rollout groups, so optimization is governed not only by the expected gradient but also by how observed groups with different empirical success count NN5 are scaled after sampling (Zheng, 28 May 2026). RL2ML treats this distinction as fundamental rather than incidental. This suggests that finite-budget training cannot be fully characterized by population-level notation such as NN6 or NN7 alone.

2. Objective family from reinforcement learning to maximum likelihood

RL2ML defines a one-parameter family of prompt-level surrogate objectives: NN8 The antiderivative is

NN9

Hence the untruncated population gradient is

xx0

This produces a continuous interpolation across three regimes (Zheng, 28 May 2026):

  • xx1: ordinary RL, since

xx2

  • xx3: maximum likelihood, since

xx4

  • xx5: “beyond-maximum-likelihood” objectives that weight low-success prompts more strongly than xx6 (Zheng, 28 May 2026).

For finite rollout budget xx7, the paper defines a rollout-aligned truncated surrogate: xx8 with population-level weight

xx9

and gradient

z∼mθ(⋅∣x)z \sim m_\theta(\cdot \mid x)0

This finite-rollout objective is the relevant one for training under a fixed rollout budget, rather than the asymptotic z∼mθ(⋅∣x)z \sim m_\theta(\cdot \mid x)1 expression alone (Zheng, 28 May 2026).

3. Exactly unbiased finite-rollout estimator

A central contribution of RL2ML is a closed-form, exactly unbiased gradient estimator for the finite-rollout truncated surrogate (Zheng, 28 May 2026). The paper first considers estimators of the form

z∼mθ(⋅∣x)z \sim m_\theta(\cdot \mid x)2

defines

z∼mθ(⋅∣x)z \sim m_\theta(\cdot \mid x)3

and uses the Bernstein basis

z∼mθ(⋅∣x)z \sim m_\theta(\cdot \mid x)4

It then shows

z∼mθ(⋅∣x)z \sim m_\theta(\cdot \mid x)5

This Bernstein representation is the mechanism by which a z∼mθ(⋅∣x)z \sim m_\theta(\cdot \mid x)6-only estimator induces a specific population weight on z∼mθ(⋅∣x)z \sim m_\theta(\cdot \mid x)7 (Zheng, 28 May 2026).

For RL2ML, the Bernstein coefficients are

z∼mθ(⋅∣x)z \sim m_\theta(\cdot \mid x)8

The corresponding group-level update scale is

z∼mθ(⋅∣x)z \sim m_\theta(\cdot \mid x)9

The estimator is then

y=f(z)y=f(z)0

and satisfies

y=f(z)y=f(z)1

The paper emphasizes that this is exact unbiasedness for the finite-rollout surrogate objective itself, not merely for an asymptotic population target (Zheng, 28 May 2026).

RL2ML also gives a control-variate form: y=f(z)y=f(z)2 with sequence-level advantage

y=f(z)y=f(z)3

At y=f(z)y=f(z)4,

y=f(z)y=f(z)5

which recovers MaxRL’s control-variate form (Zheng, 28 May 2026).

4. Group-level update geometry and the y=f(z)y=f(z)6 transition

RL2ML distinguishes population weighting from group-level update geometry by focusing on the empirical success count y=f(z)y=f(z)7 of a rollout group (Zheng, 28 May 2026). The stochastic update can be written as

y=f(z)y=f(z)8

so y=f(z)y=f(z)9 directly controls the expected update magnitude and within-group variance conditional on the observed group.

For RL2ML, the conditional moments are

r(x,z):=I{f(z)=y∗(x)}.r(x,z) := I\{f(z)=y^*(x)\}.0

and

r(x,z):=I{f(z)=y∗(x)}.r(x,z) := I\{f(z)=y^*(x)\}.1

Thus r(x,z):=I{f(z)=y∗(x)}.r(x,z) := I\{f(z)=y^*(x)\}.2 is the operative quantity that determines how observed groups are amplified or attenuated (Zheng, 28 May 2026).

The paper derives

r(x,z):=I{f(z)=y∗(x)}.r(x,z) := I\{f(z)=y^*(x)\}.3

which reveals a subcritical-supercritical update-scale transition at

r(x,z):=I{f(z)=y∗(x)}.r(x,z) := I\{f(z)=y^*(x)\}.4

This yields four regimes (Zheng, 28 May 2026):

  • RL endpoint (r(x,z):=I{f(z)=y∗(x)}.r(x,z) := I\{f(z)=y^*(x)\}.5):

r(x,z):=I{f(z)=y∗(x)}.r(x,z) := I\{f(z)=y^*(x)\}.6

  • Subcritical regime (r(x,z):=I{f(z)=y∗(x)}.r(x,z) := I\{f(z)=y^*(x)\}.7): r(x,z):=I{f(z)=y∗(x)}.r(x,z) := I\{f(z)=y^*(x)\}.8 is increasing in r(x,z):=I{f(z)=y∗(x)}.r(x,z) := I\{f(z)=y^*(x)\}.9, and

xx0

  • ML boundary / MaxRL boundary (xx1):

xx2

  • Supercritical regime (xx3): xx4 is decreasing in xx5, and

xx6

This distinction is one of RL2ML’s main conceptual claims: xx7 is a structural boundary in the sample-level geometry of finite-rollout optimization, not merely the point corresponding to xx8 in population notation (Zheng, 28 May 2026).

5. Calibrated metric gain and exact variance decomposition

RL2ML argues that the best surrogate objective is determined neither by proximity to maximum likelihood nor by the population-level weight alone (Zheng, 28 May 2026). Instead, it depends jointly on the evaluation metric, local prompt sensitivity, and estimator variance.

For a prompt-separable validation metric

xx9

the paper defines

pθ(x):=Ez∼mθ(⋅∣x)[r(x,z)].p_\theta(x) := E_{z\sim m_\theta(\cdot\mid x)}[r(x,z)].0

pθ(x):=Ez∼mθ(⋅∣x)[r(x,z)].p_\theta(x) := E_{z\sim m_\theta(\cdot\mid x)}[r(x,z)].1

After calibrating all candidate objectives to the same update length pθ(x):=Ez∼mθ(⋅∣x)[r(x,z)].p_\theta(x) := E_{z\sim m_\theta(\cdot\mid x)}[r(x,z)].2,

pθ(x):=Ez∼mθ(⋅∣x)[r(x,z)].p_\theta(x) := E_{z\sim m_\theta(\cdot\mid x)}[r(x,z)].3

the first-order metric-gain criterion is

pθ(x):=Ez∼mθ(⋅∣x)[r(x,z)].p_\theta(x) := E_{z\sim m_\theta(\cdot\mid x)}[r(x,z)].4

The local gain is

pθ(x):=Ez∼mθ(⋅∣x)[r(x,z)].p_\theta(x) := E_{z\sim m_\theta(\cdot\mid x)}[r(x,z)].5

This formalizes the claim that the preferred surrogate under a fixed rollout budget must be evaluated relative to a concrete downstream metric, not only relative to ML-like weighting (Zheng, 28 May 2026).

The variance analysis is equally explicit. With

pθ(x):=Ez∼mθ(⋅∣x)[r(x,z)].p_\theta(x) := E_{z\sim m_\theta(\cdot\mid x)}[r(x,z)].6

the paper proves

pθ(x):=Ez∼mθ(⋅∣x)[r(x,z)].p_\theta(x) := E_{z\sim m_\theta(\cdot\mid x)}[r(x,z)].7

The corresponding conditional mean-squared deviation is

pθ(x):=Ez∼mθ(⋅∣x)[r(x,z)].p_\theta(x) := E_{z\sim m_\theta(\cdot\mid x)}[r(x,z)].8

The two terms separate count variance from within-success variance (Zheng, 28 May 2026). A plausible implication is that supercritical objectives can improve hard-prompt emphasis while simultaneously incurring a distinct finite-rollout noise cost.

The remaining degree of freedom is therefore cast as a one-dimensional optimization problem: pθ(x):=Ez∼mθ(⋅∣x)[r(x,z)].p_\theta(x) := E_{z\sim m_\theta(\cdot\mid x)}[r(x,z)].9 where

S(x,z):=∇θlog⁡mθ(z∣x).S(x,z):=\nabla_\theta\log m_\theta(z\mid x).0

This is presented as a principled alternative to treating objective choice as an unconstrained hyperparameter search (Zheng, 28 May 2026).

6. Position within the broader RL-for-ML landscape

RL2ML is primarily a theoretical and analytical contribution rather than a benchmark-heavy empirical paper (Zheng, 28 May 2026). Its experiments are conceptual figures, stylized finite-horizon selection plots, and implementation guidance for Verl-style RLVR systems, rather than large real-task comparisons (Zheng, 28 May 2026). The paper’s main value is to reinterpret objective design in finite-rollout RLVR through estimator-objective alignment, group-level success-count geometry, and variance-aware metric optimization (Zheng, 28 May 2026).

In a broader methodological sense, RL2ML belongs to a line of work in which reinforcement learning is used to shape ML systems rather than only downstream task outputs. That broader pattern appears in several distinct settings:

Work RL target Core role of RL
RL2ML finite-rollout surrogate objective design aligns estimator and objective under fixed rollout budget
ReMix discrete LoRA routing trains a non-differentiable router with REINFORCE/RLOO (Qiu et al., 10 Mar 2026)
R2-Reasoner task decomposition and model allocation optimizes a multi-model routing policy with GRPO (Shao et al., 6 Jun 2025)
Retrv-R1 retrieval reasoning and inspection actions trains correctness-efficiency retrieval behavior with GRPO (Zhu et al., 3 Oct 2025)

These neighboring examples differ in mechanism, but they share a structural feature with RL2ML: reinforcement learning is used to optimize architectural, routing, or systems-level decisions rather than only a final response distribution. This suggests that RL2ML’s emphasis on finite-sample update geometry is relevant beyond correctness-based RLVR, especially in settings where rollout groups, sparse rewards, and budgeted stochastic updates are central.

RL2ML’s main technical takeaway is that finite-rollout objective design should not be reduced to a binary choice between RL and maximum likelihood. Under a fixed rollout budget, the effective training signal depends on the surrogate family

S(x,z):=∇θlog⁡mθ(z∣x).S(x,z):=\nabla_\theta\log m_\theta(z\mid x).1

the observed success count S(x,z):=∇θlog⁡mθ(z∣x).S(x,z):=\nabla_\theta\log m_\theta(z\mid x).2, the induced group-level scale S(x,z):=∇θlog⁡mθ(z∣x).S(x,z):=\nabla_\theta\log m_\theta(z\mid x).3, the target evaluation metric, and the resulting estimator variance (Zheng, 28 May 2026). Within that framework, S(x,z):=∇θlog⁡mθ(z∣x).S(x,z):=\nabla_\theta\log m_\theta(z\mid x).4 remains important as the maximum-likelihood boundary and the point where S(x,z):=∇θlog⁡mθ(z∣x).S(x,z):=\nabla_\theta\log m_\theta(z\mid x).5 becomes constant across successful groups, but it is not presented as a universal optimum (Zheng, 28 May 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to RL2ML.