Papers
Topics
Authors
Recent
Search
2000 character limit reached

Parameter-Level Characterization of RLVR

Updated 12 November 2025
  • Parameter-Level Characterization of RLVR is a detailed analysis that mathematically and empirically explains how reinforcement learning with verifiable rewards updates model weights.
  • The study shows that KL constraints, spectral geometry, and floating-point precision collectively steer updates into low-curvature, off-principal subspaces, preserving pretrained behavior.
  • Empirical comparisons reveal that RLVR achieves minimal spectral drift and reduced catastrophic forgetting relative to supervised fine-tuning, informing optimized fine-tuning strategies.

Reinforcement Learning with Verifiable Rewards (RLVR) is a paradigm for fine-tuning LLMs, primarily leveraging outcome-based, binary-verifiable supervision to improve model reasoning on tasks such as mathematics and programming. Parameter-level characterization of RLVR concerns the mathematical and empirical description of how the weights of these models evolve during RLVR post-training, addressing both the directionality and geometry of updates, their optimization regimes, and the factors affecting convergence, stability, and sparsity. This analytic perspective both distinguishes RLVR from supervised fine-tuning (SFT) and exposes the mechanistic underpinnings of its unique learning dynamics.

1. Parameter-Level Update Formulation in RLVR

RLVR optimizes model parameters θ to maximize expected reward via reinforcement learning, typically subject to a KL constraint relative to a reference policy. The one-step RLVR update can be written as the solution to

θ+∈arg⁡min⁡θDKL(q~β ∥ πθ)\theta^+ \in \arg\min_\theta D_{\mathrm{KL}}(\tilde q_\beta\,\|\,\pi_\theta)

where

q~β(y∣x)∝πref(y∣x) eR(x,y)/β\tilde q_\beta(y|x)\propto\pi_{\rm ref}(y|x)\,e^{R(x,y)/\beta}

and R(x,y)R(x,y) is a verifiable reward.

At the parameter block level (e.g., for a linear layer WW), curvature-constrained update size is bounded as

∥ΔW∥F≤2Kμ\|\Delta W\|_F \leq \sqrt{\tfrac{2K}{\mu}}

where K≡DKL(πθ+∥πθ)K \equiv D_{\mathrm{KL}}(\pi_{\theta^+}\|\pi_\theta) and μ\mu is the smallest eigenvalue of the local Fisher information submatrix.

Within RLVR frameworks such as PACS, the policy model parameterizes both the policy πθ\pi_\theta and a score function sθ(x,y)s_\theta(x,y), mapping to a probability of correctness. PACS reformulates RLVR’s objective via a supervised cross-entropy loss over (x,y,r(y))(x, y, r(y)) triples:

q~β(y∣x)∝πref(y∣x) eR(x,y)/β\tilde q_\beta(y|x)\propto\pi_{\rm ref}(y|x)\,e^{R(x,y)/\beta}0

with the gradient decomposition

q~β(y∣x)∝πref(y∣x) eR(x,y)/β\tilde q_\beta(y|x)\propto\pi_{\rm ref}(y|x)\,e^{R(x,y)/\beta}1

resulting in an implicit actor–critic structure operating within a single parameter set.

2. Three-Gate Theory of RLVR Parameter Dynamics

The evolution of parameters under RLVR is explained mechanistically by the Three-Gate Theory:

  1. KL Anchor (Gate I): Each policy update is strictly KL-constrained, bounding weight change in parameter space. For sufficiently small steps,

q~β(y∣x)∝πref(y∣x) eR(x,y)/β\tilde q_\beta(y|x)\propto\pi_{\rm ref}(y|x)\,e^{R(x,y)/\beta}2

which in turn controls the Fisher-weighted q~β(y∣x)∝πref(y∣x) eR(x,y)/β\tilde q_\beta(y|x)\propto\pi_{\rm ref}(y|x)\,e^{R(x,y)/\beta}3 norm of q~β(y∣x)∝πref(y∣x) eR(x,y)/β\tilde q_\beta(y|x)\propto\pi_{\rm ref}(y|x)\,e^{R(x,y)/\beta}4.

  1. Model Geometry (Gate II): The geometry of the pretrained model, embodied in its spectral decomposition, dictates that—under the KL leash—updates preferentially avoid principal (high-curvature) subspaces. Specifically, for a pretrained block q~β(y∣x)∝πref(y∣x) eR(x,y)/β\tilde q_\beta(y|x)\propto\pi_{\rm ref}(y|x)\,e^{R(x,y)/\beta}5:

q~β(y∣x)∝πref(y∣x) eR(x,y)/β\tilde q_\beta(y|x)\propto\pi_{\rm ref}(y|x)\,e^{R(x,y)/\beta}6

Wedin’s sin–Θ theorem and Weyl’s bounds guarantee that for sufficiently small q~β(y∣x)∝πref(y∣x) eR(x,y)/β\tilde q_\beta(y|x)\propto\pi_{\rm ref}(y|x)\,e^{R(x,y)/\beta}7, subspace rotation (q~β(y∣x)∝πref(y∣x) eR(x,y)/β\tilde q_\beta(y|x)\propto\pi_{\rm ref}(y|x)\,e^{R(x,y)/\beta}8) and spectral drift (q~β(y∣x)∝πref(y∣x) eR(x,y)/β\tilde q_\beta(y|x)\propto\pi_{\rm ref}(y|x)\,e^{R(x,y)/\beta}9) are minimal, keeping updates in spectrum-preserving, low-curvature directions.

  1. Precision (Gate III): Given the floating-point format (e.g., bfloat16), micro-updates below a threshold (relative ULP R(x,y)R(x,y)0 0.2–0.4%) are “invisible,” causing updates in non-preferred subspaces to appear sparse, although energy is distributed predominantly off-principal.

These gates cooperate to steer RLVR updates into subspaces unlikely to disrupt core pretrained behavior, resulting in improved stability and reduced catastrophic forgetting.

3. Metrics for Parameter-Space Characterization

Empirical analysis of RLVR’s parameter evolution employs quantitative metrics:

Metric Definition Significance
Spectral Drift (R(x,y)R(x,y)1) R(x,y)R(x,y)2 Captures magnitude of singular value shift post-update
Principal-Subspace Rotation (R(x,y)R(x,y)3) R(x,y)R(x,y)4 Measures maximum angle rotation of top-R(x,y)R(x,y)5 invariant subspaces
Off-Principal Update Alignment (R(x,y)R(x,y)6) R(x,y)R(x,y)7 Proportion of update norm lying outside top-R(x,y)R(x,y)8 directions

RLVR consistently exhibits R(x,y)R(x,y)9 and subspace rotation WW0 compared to SFT’s drift WW1 and rotations WW2 on models such as DS-Qwen-1.5B and Qwen3-8B.

4. Gradient Gap, Alignment, and Convergence Thresholds

A central theoretical quantity for RLVR optimization is the parameter-level (token-level) Gradient Gap. For an autoregressive policy WW3 and a fixed prompt WW4,

WW5

where WW6 and WW7 are the conditional distributions over successful and failing trajectories. At iteration WW8, the update direction WW9 yields gap alignment ∥ΔW∥F≤2Kμ\|\Delta W\|_F \leq \sqrt{\tfrac{2K}{\mu}}0.

Convergence of the RLVR procedure is controlled by the magnitude and alignment of ∥ΔW∥F≤2Kμ\|\Delta W\|_F \leq \sqrt{\tfrac{2K}{\mu}}1 and the learning rate ∥ΔW∥F≤2Kμ\|\Delta W\|_F \leq \sqrt{\tfrac{2K}{\mu}}2, with the main update bound:

∥ΔW∥F≤2Kμ\|\Delta W\|_F \leq \sqrt{\tfrac{2K}{\mu}}3

where ∥ΔW∥F≤2Kμ\|\Delta W\|_F \leq \sqrt{\tfrac{2K}{\mu}}4 incorporates response length ∥ΔW∥F≤2Kμ\|\Delta W\|_F \leq \sqrt{\tfrac{2K}{\mu}}5 and token score norms. A critical threshold emerges: ∥ΔW∥F≤2Kμ\|\Delta W\|_F \leq \sqrt{\tfrac{2K}{\mu}}6 requiring the learning rate to shrink with increasing response length ∥ΔW∥F≤2Kμ\|\Delta W\|_F \leq \sqrt{\tfrac{2K}{\mu}}7 and as success probability ∥ΔW∥F≤2Kμ\|\Delta W\|_F \leq \sqrt{\tfrac{2K}{\mu}}8. A fixed ∥ΔW∥F≤2Kμ\|\Delta W\|_F \leq \sqrt{\tfrac{2K}{\mu}}9 will eventually violate this constraint, resulting in stagnation below perfect accuracy.

This theory mathematically justifies practical heuristics like length normalization, scaling K≡DKL(πθ+∥πθ)K \equiv D_{\mathrm{KL}}(\pi_{\theta^+}\|\pi_\theta)0, and performance-aware scaling K≡DKL(πθ+∥πθ)K \equiv D_{\mathrm{KL}}(\pi_{\theta^+}\|\pi_\theta)1, and explains the limitations of naïve fixed-rate or step-size approaches.

5. Empirical Findings and Comparisons with Supervised Fine-Tuning

Empirical studies substantiating the above theory show:

  • Spectrum Preservation: RLVR’s trajectory maintains spectral drift and subspace rotation substantially below SFT, indicating preservation of core model representations.
  • Off-Principal Alignment: RLVR updates are consistently aligned with low-curvature, non-principal directions (K≡DKL(πθ+∥πθ)K \equiv D_{\mathrm{KL}}(\pi_{\theta^+}\|\pi_\theta)2), and overlap with principal weight masks falls below random, confirming avoidance of the “principal” spectrum.
  • Invariance and Consistency: The parameter regions updated under RLVR are highly consistent across RNG seeds, datasets, and RL variants, indicating strong model-intrinsic bias.
  • Sparsity Artifact: The apparent sparsity of RLVR-induced parameter changes is explained by micro-updates masked by bfloat16 ULP (Gate III), not true inactivity.
  • Intervention Validation: Destructive interventions (rotating principal subspaces, head permutations) collapse RLVR’s parameter-localization effect, demonstrating causality for Gate II.

Comparatively, SFT distorts principal directions, increases spectral drift, and tends to induce greater catastrophic forgetting. RLVR achieves equivalent or superior reasoning improvements with milder parameter adjustments, emphasizing its regime of minimal-disruption fine-tuning.

6. Implications for Fine-Tuning Strategies and Architecture Design

The parameter-level characterization of RLVR challenges the direct adaptation of SFT-era fine-tuning heuristics and PEFT methods. For instance, LoRA targeting low-rank, spectrum-complement subspaces matches full RLVR post-training dynamics, while PiSSA or masking based on principal components impedes or collapses RLVR effectiveness. Freezing principal subspaces slows learning and accuracy, whereas freezing the complement preserves RLVR trajectory, further emphasizing the protocol’s off-principal orientation.

A plausible implication is that new PEFT strategies for RLVR should explicitly exploit this geometric and spectral bias (“geometry-aware, RLVR-native learning algorithms”), moving away from principal-aligned SFT heuristics.

7. Summary and Outlook

Parameter-level analysis of RLVR establishes that policy optimization is strongly KL-anchored (Gate I), geometry-preserving (Gate II), and arguably ULP-masked for micro-updates (Gate III), resulting in off-principal, low-curvature parameter trajectory. Theoretical constructs such as the Gradient Gap, along with sharp step-size constraints, mathematically explain convergence behavior, stability advantages, and deployment heuristics unique to RLVR. Empirical results confirm that RLVR delivers robust, minimally disruptive post-training, contrasting starkly with SFT. These findings provide a substrate for further refinement of RLVR methods and optimization-aware model adaptation strategies (Li et al., 2 Sep 2025, Zhu et al., 11 Nov 2025, Suk et al., 9 Oct 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Parameter-Level Characterization of RLVR.