---
title: 'RL2ML: Finite-Rollout Surrogate Objectives'
url: https://www.emergentmind.com/topics/rl2ml
type: topic
---

# RL2ML: Finite-Rollout Surrogate Objectives

RL2ML denotes a family of **finite-rollout surrogate objectives** for correctness-based **reinforcement learning with verifiable rewards** (RLVR), introduced to separate the **population-level objective optimized in expectation** from the **stochastic update geometry induced by finite rollout groups** [2605.30154]. In this formulation, language-model training from binary correctness signals is not treated as a single fixed point between ordinary reinforcement learning and maximum likelihood, but as a continuous one-parameter family indexed by \(\gamma\), with exact finite-rollout estimators aligned to the corresponding surrogate objective [2605.30154]. The framework is motivated by the observation that in RLVR, the objective optimized in expectation and the updates produced under a fixed rollout budget \(N\) are often conflated, even though finite-group success counts can induce qualitatively different optimization behavior [2605.30154].

## 1. RL2ML problem formulation

RL2ML is defined in the setting of **correctness-based RL with verifiable rewards**. For each prompt \(x\), a model samples a latent rollout \(z \sim m_\theta(\cdot \mid x)\), applies deterministic decoding \(y=f(z)\), and receives binary reward
\[
r(x,z) := I\{f(z)=y^*(x)\}.
\]
For a fixed prompt \(x\), the model success probability is
\[
p_\theta(x) := E_{z\sim m_\theta(\cdot\mid x)}[r(x,z)].
\]
The score function is
\[
S(x,z):=\nabla_\theta\log m_\theta(z\mid x).
\]
If \(N\) rollouts are sampled independently, the paper defines
\[
r_i:=r(x,z_i), \qquad S_i:=S(x,z_i), \qquad K:=\sum_{i=1}^N r_i,
\]
and, when \(K\ge 1\),
\[
\bar S_K(x) := \frac1K\sum_{i=1}^N r_iS_i.
\]
It further introduces
\[
\mu_x:=E[S\mid r=1,x],\qquad \Sigma_x:=Cov(S\mid r=1,x),
\]
with the identities
\[
\mu_x=\nabla_\theta\log p_\theta(x),\qquad \nabla_\theta p_\theta(x)=p_\theta(x)\mu_x
\]
[2605.30154].

The central problem is that RLVR training uses **finite rollout groups**, so optimization is governed not only by the expected gradient but also by how observed groups with different empirical success count \(K\) are scaled after sampling [2605.30154]. RL2ML treats this distinction as fundamental rather than incidental. This suggests that finite-budget training cannot be fully characterized by population-level notation such as \(\nabla_\theta p\) or \(\nabla_\theta \log p\) alone.

## 2. Objective family from reinforcement learning to maximum likelihood

RL2ML defines a one-parameter family of prompt-level surrogate objectives:
\[
\mathcal{J}^{\mathrm{RL2ML}_\gamma(x)} = \phi_\gamma(p_\theta(x)),\ \text{where}\ \phi_\gamma'(p)=p^{-\gamma}.
\]
The antiderivative is
\[
\phi_\gamma(p) = \begin{cases} \dfrac{p^{1-\gamma}-1}{1-\gamma}, & \gamma\neq 1,\\[0.8em] \log p, & \gamma=1. \end{cases}
\]
Hence the untruncated population gradient is
\[
\nabla_\theta\mathcal{J}^{\mathrm{RL2ML}_\gamma(x)} = \bigl(p_\theta(x)\bigr)^{-\gamma}\nabla_\theta p_\theta(x).
\]
This produces a continuous interpolation across three regimes [2605.30154]:

- **\(\gamma=0\)**: ordinary RL, since
  \[
  \nabla_\theta\mathcal{J}^{\mathrm{RL2ML}_0(x)}=\nabla_\theta p.
  \]

- **\(\gamma=1\)**: maximum likelihood, since
  \[
  \nabla_\theta\mathcal{J}^{\mathrm{RL2ML}_1(x)}=\nabla_\theta\log p.
  \]

- **\(\gamma>1\)**: “beyond-maximum-likelihood” objectives that weight low-success prompts more strongly than \(1/p\) [2605.30154].

For finite rollout budget \(N\), the paper defines a rollout-aligned truncated surrogate:
\[
\mathcal{J}^{\mathrm{RL2ML}_{\gamma,N}(x)} := \sum_{k=1}^N \frac{(\gamma)_{k-1}}{(k-1)!\,k}pass@k(x),
\]
with population-level weight
\[
w^{\mathrm{RL2ML}_{\gamma,N}}(p) := \sum_{m=0}^{N-1}\frac{(\gamma)_m}{m!}(1-p)^m,
\]
and gradient
\[
\nabla_\theta\mathcal{J}^{\mathrm{RL2ML}_{\gamma,N}(x)} = w^{\mathrm{RL2ML}_{\gamma,N}}(p_\theta(x))\nabla_\theta p_\theta(x).
\]
This finite-rollout objective is the relevant one for training under a fixed rollout budget, rather than the asymptotic \(p^{-\gamma}\nabla_\theta p\) expression alone [2605.30154].

## 3. Exactly unbiased finite-rollout estimator

A central contribution of RL2ML is a **closed-form, exactly unbiased gradient estimator** for the finite-rollout truncated surrogate [2605.30154]. The paper first considers estimators of the form
\[
\hat g_f(x)=f(K)\sum_{i=1}^N r_iS_i,
\]
defines
\[
\beta_m:=Nf(m+1),\ \text{for}\ m=0,\ldots,N-1,
\]
and uses the Bernstein basis
\[
B_{m,N-1}(p):=\binom{N-1}{m}p^m(1-p)^{N-1-m}.
\]
It then shows
\[
E[\hat g_f(x)\mid x] = \left(\sum_{m=0}^{N-1}\beta_mB_{m,N-1}(p)\right)\nabla_\theta p.
\]
This Bernstein representation is the mechanism by which a \(K\)-only estimator induces a specific population weight on \(\nabla_\theta p\) [2605.30154].

For RL2ML, the Bernstein coefficients are
\[
\beta_{K-1}^{(\gamma,N)} = \frac{\Gamma(N+\gamma)}{\Gamma(N)} \frac{\Gamma(K)}{\Gamma(K+\gamma)},\ \text{for}\ K=1,\ldots,N.
\]
The corresponding group-level update scale is
\[
\alpha_K^{(\gamma,N)} = \frac{K}{N}\beta_{K-1}^{(\gamma,N)} = \frac{\Gamma(N+\gamma)}{\Gamma(N+1)} \frac{\Gamma(K+1)}{\Gamma(K+\gamma)}.
\]
The estimator is then
\[
\hat g_{\gamma,N}^{\mathrm{RL2ML}(x)} = \alpha_K^{(\gamma,N)}\bar S_K,
\]
and satisfies
\[
E[\hat g_{\gamma,N}^{\mathrm{RL2ML}(x)}\mid x] = \nabla_\theta\mathcal{J}^{\mathrm{RL2ML}_{\gamma,N}(x)}.
\]
The paper emphasizes that this is **exact unbiasedness for the finite-rollout surrogate objective itself**, not merely for an asymptotic population target [2605.30154].

RL2ML also gives a control-variate form:
\[
\tilde g_{\gamma,N}(x) = \frac1N\sum_{i=1}^N\bigl(\beta_{K-1}^{(\gamma,N)}r_i-1\bigr)S_i,
\]
with sequence-level advantage
\[
A_i^{(\gamma,N)}=\beta_{K-1}^{(\gamma,N)}r_i-1.
\]
At \(\gamma=1\),
\[
\beta_{K-1}^{(1,N)}=\frac{N}{K},\qquad A_i^{(1,N)}=\frac{N}{K}r_i-1,
\]
which recovers MaxRL’s control-variate form [2605.30154].

## 4. Group-level update geometry and the \(\gamma=1\) transition

RL2ML distinguishes **population weighting** from **group-level update geometry** by focusing on the empirical success count \(K\) of a rollout group [2605.30154]. The stochastic update can be written as
\[
\hat g_f(x) = \alpha_K\bar S_K,
\]
so \(\alpha_K\) directly controls the expected update magnitude and within-group variance conditional on the observed group.

For RL2ML, the conditional moments are
\[
E[\hat g_f(x)\mid x,K=k]=\alpha_k\mu_x,
\]
and
\[
Cov(\hat g_f(x)\mid x,K=k)=\frac{\alpha_k^2}{k}\Sigma_x.
\]
Thus \(\alpha_K\) is the operative quantity that determines how observed groups are amplified or attenuated [2605.30154].

The paper derives
\[
\frac{\alpha_{K+1}^{(\gamma,N)}}{\alpha_K^{(\gamma,N)}} = \frac{K+1}{K+\gamma},
\]
which reveals a **subcritical-supercritical update-scale transition** at
\[
\gamma = 1.
\]
This yields four regimes [2605.30154]:

- **RL endpoint** (\(\gamma=0\)):
  \[
  \alpha_K^{(0,N)}=K/N.
  \]

- **Subcritical regime** (\(0<\gamma<1\)): \(\alpha_K^{(\gamma,N)}\) is increasing in \(K\), and
  \[
  K/N<\alpha_K^{(\gamma,N)}<1,\qquad 1\le K<N.
  \]

- **ML boundary / MaxRL boundary** (\(\gamma=1\)):
  \[
  \alpha_K^{(1,N)}\equiv 1.
  \]

- **Supercritical regime** (\(\gamma>1\)): \(\alpha_K^{(\gamma,N)}\) is decreasing in \(K\), and
  \[
  \alpha_K^{(\gamma,N)}>1,\qquad \text{for all } K<N.
  \]

This distinction is one of RL2ML’s main conceptual claims: \(\gamma=1\) is a structural boundary in the **sample-level** geometry of finite-rollout optimization, not merely the point corresponding to \(\log p\) in population notation [2605.30154].

## 5. Calibrated metric gain and exact variance decomposition

RL2ML argues that the best surrogate objective is determined neither by proximity to maximum likelihood nor by the population-level weight alone [2605.30154]. Instead, it depends jointly on the evaluation metric, local prompt sensitivity, and estimator variance.

For a prompt-separable validation metric
\[
V_{\mathcal C}(\theta) = \sum_{x\in\mathcal C}v_x(p_x),
\]
the paper defines
\[
\ell_x:=\|\nabla_{\theta_x}p_x\|_2^2,
\]
\[
A(\gamma):=\sum_{x\in\mathcal C}v_x'(p_x)\ell_x w_{\gamma,N}(p_x),
\qquad
B(\gamma):=\sum_{x\in\mathcal C}\ell_x w_{\gamma,N}(p_x)^2.
\]
After calibrating all candidate objectives to the same update length \(c\),
\[
\theta_\gamma^+ := \theta+\eta_\gamma D_{\gamma,\mathcal C}(\theta),\qquad
\eta_\gamma:=\frac{c}{\sqrt{B(\gamma)}},
\]
the first-order metric-gain criterion is
\[
U(\gamma) = \frac{A(\gamma)}{\sqrt{B(\gamma)}}
= \frac{ \sum_{x\in\mathcal C} v_x'(p_x)\,w_{\gamma,N}(p_x)\,\ell_x }
{ \sqrt{ \sum_{x\in\mathcal C} w_{\gamma,N}(p_x)^2\,\ell_x } }.
\]
The local gain is
\[
V_{\mathcal C}(\theta_\gamma^+)-V_{\mathcal C}(\theta) = c\,U(\gamma)+\mathcal{O}(c^2).
\]
This formalizes the claim that the preferred surrogate under a fixed rollout budget must be evaluated relative to a concrete downstream metric, not only relative to ML-like weighting [2605.30154].

The variance analysis is equally explicit. With
\[
a_K(\gamma):=\alpha_K^{(\gamma,N)}I\{K\ge 1\},
\]
the paper proves
\[
Cov(\hat g_{\gamma,N}(x)\mid x) =
Var(a_K(\gamma)\mid x)\mu_x\mu_x^\top +
E\!\left[\frac{a_K(\gamma)^2}{K}\middle|x\right]\Sigma_x.
\]
The corresponding conditional mean-squared deviation is
\[
E\!\left[\|\hat g_{\gamma,N}(x)-E[\hat g_{\gamma,N}(x)\mid x]\|_2^2\middle|x\right]
=
Var(a_K(\gamma)\mid x)\|\mu_x\|_2^2 +
E\!\left[\frac{a_K(\gamma)^2}{K}\middle|x\right]tr(\Sigma_x).
\]
The two terms separate **count variance** from **within-success variance** [2605.30154]. A plausible implication is that supercritical objectives can improve hard-prompt emphasis while simultaneously incurring a distinct finite-rollout noise cost.

The remaining degree of freedom is therefore cast as a one-dimensional optimization problem:
\[
\gamma_\lambda^* =
\operatorname*{arg\,max}_{\gamma\in\Gamma}
\left[U(\gamma)-\lambda_{\mathrm{var}}R(\gamma)^{1/2}\right],
\]
where
\[
R(\gamma) := \sum_{x\in\mathcal C}
\left[
Var(a_K(\gamma)\mid x)\|\mu_x\|_2^2 +
E\!\left[\frac{a_K(\gamma)^2}{K}\middle|x\right]tr(\Sigma_x)
\right].
\]
This is presented as a principled alternative to treating objective choice as an unconstrained hyperparameter search [2605.30154].

## 6. Position within the broader RL-for-ML landscape

RL2ML is primarily a **theoretical and analytical** contribution rather than a benchmark-heavy empirical paper [2605.30154]. Its experiments are conceptual figures, stylized finite-horizon selection plots, and implementation guidance for Verl-style RLVR systems, rather than large real-task comparisons [2605.30154]. The paper’s main value is to reinterpret objective design in finite-rollout RLVR through estimator-objective alignment, group-level success-count geometry, and variance-aware metric optimization [2605.30154].

In a broader methodological sense, RL2ML belongs to a line of work in which reinforcement learning is used to shape ML systems rather than only downstream task outputs. That broader pattern appears in several distinct settings:

| Work | RL target | Core role of RL |
|---|---|---|
| RL2ML | finite-rollout surrogate objective design | aligns estimator and objective under fixed rollout budget |
| ReMix | discrete LoRA routing | trains a non-differentiable router with REINFORCE/RLOO [2603.10160] |
| R2-Reasoner | task decomposition and model allocation | optimizes a multi-model routing policy with GRPO [2506.05901] |
| Retrv-R1 | retrieval reasoning and inspection actions | trains correctness-efficiency retrieval behavior with GRPO [2510.02745] |

These neighboring examples differ in mechanism, but they share a structural feature with RL2ML: reinforcement learning is used to optimize **architectural, routing, or systems-level decisions** rather than only a final response distribution. This suggests that RL2ML’s emphasis on finite-sample update geometry is relevant beyond correctness-based RLVR, especially in settings where rollout groups, sparse rewards, and budgeted stochastic updates are central.

RL2ML’s main technical takeaway is that finite-rollout objective design should not be reduced to a binary choice between RL and maximum likelihood. Under a fixed rollout budget, the effective training signal depends on the surrogate family
\[
\mathcal{J}^{\mathrm{RL2ML}_{\gamma,N}},
\]
the observed success count \(K\), the induced group-level scale \(\alpha_K^{(\gamma,N)}\), the target evaluation metric, and the resulting estimator variance [2605.30154]. Within that framework, \(\gamma=1\) remains important as the maximum-likelihood boundary and the point where \(\alpha_K\) becomes constant across successful groups, but it is not presented as a universal optimum [2605.30154].

Source: https://www.emergentmind.com/topics/rl2ml