Papers
Topics
Authors
Recent
Search
2000 character limit reached

Score Centering Stabilizes Off-policy Reinforcement Learning

Published 17 Sep 2026 in cs.LG | (2609.20807v1)

Abstract: Reinforcement learning (RL) of LLMs is notoriously sensitive to small differences between training and inference engines, often referred to as the training-inference mismatch (TIM). However, completely eliminating TIM is impractical, as it would come at a major cost to rollout efficiency. In this paper, we show that the instability of RL under TIM is primarily caused by drift: a persistent bias between training and inference engines that accumulates with every training step. We derive an additive "score centering" correction term that stabilizes RL under TIM by canceling drift. When training models from 0.6B to 30B parameters, score centering alone matches or outperforms methods based on importance sampling under quantization, with the gap growing as the mismatch becomes more severe. Because the correction is additive, score centering also composes with importance sampling -- their composition outperforms pure importance-sampling baselines in our staleness experiments.

Authors (2)

Summary

  • The paper introduces score centering, a method that eliminates the drift term in policy-gradient estimators, stabilizing off-policy reinforcement learning during training-inference mismatch (TIM).
  • Score centering significantly outperforms importance sampling and other baselines under severe quantization and quantized KB-cache settings.
  • The methodology is computationally efficient, using top-k approximation to avoid storing full vocabularies while demonstrating comparable performance to full-vocabulary implementations, making it scalable for practical applications.

Problem setting and central claim

“Score Centering Stabilizes Off-policy Reinforcement Learning” (2609.20807) studies the training-inference mismatch (TIM) that arises when LLM rollouts are generated by an inference engine with policy qθq_\theta, while gradients are computed by a training engine representing policy pθp_\theta. Although the two engines are nominally instances of the same model, differences in numerical precision, kernel execution, batching, quantization, code paths, routing, and checkpoint freshness can make qθ≠pθq_\theta \ne p_\theta.

The paper’s central claim is that the principal source of instability under TIM is not mismatch per se, but an accumulating drift term in the policy-gradient estimator. This drift causes the trainer to distill toward the sampler, even when the reward contains no discriminative information. Because the sampler is subsequently refreshed from the trainer, the resulting feedback loop compounds the bias over training. The authors derive an additive correction, score centering (SC), that cancels this drift exactly at each prefix without using importance ratios. Across Qwen3 models from 0.6B to 30B parameters, SC is competitive with or superior to importance-sampling corrections under severe quantization and staleness.

Training-inference mismatch and the source of instability

For an on-policy policy-gradient method, the expected update is the reward-weighted score,

Epθ[R∇θlog⁡pθ(y)].\mathbb{E}_{p_\theta}\left[R \nabla_\theta \log p_\theta(y)\right].

When rollouts are sampled from qθq_\theta but scored under pθp_\theta, this identity no longer directly applies. Exact correction through importance sampling is possible, but the likelihood ratio pθ(y)/qθ(y)p_\theta(y)/q_\theta(y) has high variance, especially for autoregressive sequences and rare tokens. Practical methods therefore clip, mask, or otherwise truncate the ratio, introducing bias.

The paper isolates a different failure mechanism by conditioning on a prefix and considering one next-token score. Let syt=∇θlog⁡pθ(yt)s_{y_t}=\nabla_\theta\log p_\theta(y_t) and let sˉ\bar{s} denote the expected score under the sampler. The off-policy update decomposes as

Eq[Rsyt]=Eq[R]sˉ+Cov⁡q(R,syt).\mathbb{E}_q[R s_{y_t}] = \mathbb{E}_q[R]\bar{s} + \operatorname{Cov}_q(R,s_{y_t}).

The covariance is the reward-dependent signal. The first term is the drift. It depends on the reward only through its conditional mean and therefore does not identify which token contributed to success. Since pθp_\theta0 is the negative gradient of the cross-entropy from the sampler distribution to the trainer distribution, this term acts as distillation toward pθp_\theta1.

On-policy, pθp_\theta2 because the expected score under pθp_\theta3 vanishes. Under TIM, however, pθp_\theta4 is generally nonzero. A sampler that is quantized, stale, or numerically inconsistent therefore becomes a moving teacher. The trainer is pushed toward that biased teacher and is then copied back to the sampler, allowing the mismatch-induced error to accumulate. This explanation distinguishes online RL from ordinary offline distillation: distillation toward a fixed teacher can converge, whereas online distillation toward a periodically refreshed and systematically biased copy can create a positive feedback loop.

The experiments also separate two instability mechanisms that are often conflated. In the authors’ Countdown experiments, offline training is stable primarily with nonnegative pθp_\theta5 rewards, whereas online training under TIM is least stable with those same rewards. Group-centered rewards containing both positive and negative values are more stable online. The result supports the paper’s assertion that offline negative-reward instability and online TIM-induced drift have different causes. The paper does not further analyze the former, which it associates with unbounded negative-log-probability updates and distribution sharpening.

Score centering

Score centering replaces each token score with

pθp_\theta6

where the expectation pθp_\theta7 is taken under the sampler’s next-token distribution. Consequently,

pθp_\theta8

and the expected update becomes only the covariance term,

pθp_\theta9

This removes the drift exactly, including under a constant reward. The remaining difference from the ideal on-policy update is that the covariance is estimated under qθ≠pθq_\theta \ne p_\theta0 rather than qθ≠pθq_\theta \ne p_\theta1. Thus SC is not an exact replacement for importance sampling in every off-policy regime; it specifically removes the additive bias caused by the nonzero expected score.

The distinction from conventional baselines and control variates is important. Reward baselines exploit the zero-mean score identity on-policy to reduce variance. SC instead subtracts the expected score itself to correct the mean of the off-policy update. Unlike importance sampling, it is additive and deterministic conditional on the prefix. It introduces neither sampled multiplicative ratios nor clipping thresholds.

SC also composes naturally with importance sampling. If an importance method assigns token weight qθ≠pθq_\theta \ne p_\theta2, SC subtracts the expected weighted score,

qθ≠pθq_\theta \ne p_\theta3

This composition preserves exact cancellation of drift for the weighted estimator. The paper therefore treats SC and importance sampling as orthogonal: SC controls the additive mismatch-induced bias, while importance weighting partially corrects the distribution under which the reward-score covariance is measured.

Efficient implementation with top-qθ≠pθq_\theta \ne p_\theta4 log probabilities

Exact SC would require the full sampler vocabulary distribution for every generated token, which is impractical for large vocabularies and long rollouts. The implementation stores only the sampler’s top-qθ≠pθq_\theta \ne p_\theta5 log probabilities and reconstructs the tail using the trainer distribution, rescaled to match the sampler’s tail mass.

For head tokens qθ≠pθq_\theta \ne p_\theta6, the approximate sampler distribution retains the sampler probabilities. For the tail, it uses qθ≠pθq_\theta \ne p_\theta7, where qθ≠pθq_\theta \ne p_\theta8 is the ratio between sampler and trainer tail mass. Because the trainer’s full expected score is zero, the tail contribution can be reduced to a correction over the top-qθ≠pθq_\theta \ne p_\theta9 tokens:

Epθ[R∇θlog⁡pθ(y)].\mathbb{E}_{p_\theta}\left[R \nabla_\theta \log p_\theta(y)\right].0

This yields a scalar loss implementable with standard autodifferentiation and stop-gradient operations. In the reported experiments, Epθ[R∇θlog⁡pθ(y)].\mathbb{E}_{p_\theta}\left[R \nabla_\theta \log p_\theta(y)\right].1 adds negligible wall-clock cost—within approximately 1% of baseline methods—and Epθ[R∇θlog⁡pθ(y)].\mathbb{E}_{p_\theta}\left[R \nabla_\theta \log p_\theta(y)\right].2 performs comparably to full-vocabulary centering.

Figure 1

Figure 1: Top-Epθ[R∇θlog⁡pθ(y)].\mathbb{E}_{p_\theta}\left[R \nabla_\theta \log p_\theta(y)\right].3 score centering with Epθ[R∇θlog⁡pθ(y)].\mathbb{E}_{p_\theta}\left[R \nabla_\theta \log p_\theta(y)\right].4 or Epθ[R∇θlog⁡pθ(y)].\mathbb{E}_{p_\theta}\left[R \nabla_\theta \log p_\theta(y)\right].5 matches full-vocabulary score centering across the tested settings.

The top-Epθ[R∇θlog⁡pθ(y)].\mathbb{E}_{p_\theta}\left[R \nabla_\theta \log p_\theta(y)\right].6 approximation remains effective even in the most severe reported 30B setting, where the sampler’s top-128 mass is less accurately concentrated because of INT4 KV-cache quantization. This result supports the practical claim that SC does not require storing prohibitively large full-vocabulary distributions. It does not, however, establish that the same tail model will remain accurate for arbitrary vocabularies, tokenizers, or substantially more pathological sampler distributions.

Experimental design

The experiments use a common REINFORCE objective with group-centered rewards and compare correction methods in isolation. This design is methodologically useful because several baselines normally bundle correction rules with other algorithmic components such as dynamic sampling, response-length penalties, or policy-version clipping. The paper instead holds the sampler, trainer, optimizer, advantage estimates, and one-step-per-batch update fixed, applying each correction to the shared objective.

The principal models and tasks are Qwen3-0.6B-Instruct on Countdown and Qwen3-30B-A3B-Base on the mathematical subset of INTELLECT-2. TIM is induced through three mechanisms: synthetic Gaussian sampler-weight perturbations, sampler quantization, and sampler staleness. The perturbations are intentionally severe to produce separation between methods within feasible compute budgets. The authors explicitly interpret these short, severe-mismatch experiments as a proxy for longer training under milder mismatch, rather than as a direct reproduction of ordinary deployment conditions.

Synthetic mismatch and drift accumulation

Under fixed Gaussian perturbations to sampler weights, larger mismatch causes earlier collapse. SC, truncated importance sampling (TIS), masked importance sampling (MIS), and their compositions are the strongest methods at lower noise levels. Under the largest perturbation, only SC and SC composed with TIS or MIS remain stable.

Figure 2

Figure 2: Increasing sampler-weight noise causes earlier collapse; SC alone or composed with TIS/MIS remains stable under the largest noise.

The timing of collapse is consistent with the drift hypothesis. For example, DPPO collapses at approximately steps 160, 80, and 20 as the noise scale increases, while TIS is stable at the smallest noise but collapses at approximately steps 180 and 40 at larger scales. The numerical pattern indicates that mismatch severity controls the rate at which harmful bias accumulates rather than merely adding a fixed perturbation to each update. This is a stronger diagnostic than comparing final performance alone: methods differ in how long they preserve stable dynamics before drift becomes dominant.

Quantization and staleness

The realistic TIM experiments use an independently quantized sampler with a BF16 trainer and a sampler refreshed only every 64 steps. Under sampler quantization, SC performs best either alone or when composed with TIS/MIS. Under staleness, the compositions dominate vanilla SC.

Figure 3

Figure 3: SC is strongest under severe sampler quantization, while SC composed with TIS or MIS is strongest under substantial update staleness.

This difference follows directly from the estimator decomposition. SC cancels the sampler-induced drift, but its residual covariance is still measured under the sampler distribution. When the trainer moves substantially during a 64-step synchronization interval, importance weighting partially corrects that distributional discrepancy, after which SC removes the remaining additive drift. The result yields a concrete operational recommendation: SC alone is particularly appropriate for numerical mismatch such as quantization, whereas SC combined with an IS correction is preferable when the dominant mismatch is checkpoint staleness.

PPO and DAPO survive the staleness condition but collapse under quantization and synthetic weight noise. The paper attributes this asymmetry to the design of their clipping regions: policy-ratio clipping addresses ratios caused by policy movement, but not necessarily systematic numerical discrepancies that do not resemble ordinary policy updates. This is a substantive claim because it challenges the assumption that standard PPO-style clipping is a general-purpose safeguard against all trainer-sampler discrepancies.

Scaling to a 30B mixture-of-experts model

The largest experiment trains Qwen3-30B-A3B-Base on INTELLECT-2 mathematics under progressively more severe sampler quantization. The sampler and trainer may also disagree in mixture-of-experts routing because router indices are not replayed, creating an additional source of TIM.

With an FP8 sampler, uncorrected policy gradient remains stable and reaches 58% training accuracy. With FP4 KV-cache quantization, vanilla policy gradient collapses within 200 steps and MIS collapses late in training, whereas SC reaches 52% and TIS reaches 51% while remaining stable. Under the more severe INT8 weight/activation and INT4 KV-cache configuration, SC reaches 30%, TIS reaches 12%, and all other methods finish below 5%.

Figure 4

Figure 4: Under progressively stronger sampler quantization, SC maintains stable training as vanilla policy gradient and several importance-sampling methods collapse.

These results provide the paper’s strongest empirical evidence for the proposed mechanism. SC does not merely match clipped importance sampling under mild mismatch; its advantage increases as mismatch becomes severe. The 30B result also indicates that the correction remains usable in a large MoE system with routing disagreement, although the limited seed counts in the most expensive experiments weaken the statistical strength of method-by-method comparisons.

Relationship between reward structure and online instability

The offline-versus-online comparison on Qwen3-1.7B provides an important qualification to the main narrative. With offline data, Epθ[R∇θlog⁡pθ(y)].\mathbb{E}_{p_\theta}\left[R \nabla_\theta \log p_\theta(y)\right].7 rewards behave like positive-sample distillation and are stable, while reward modes containing negative values can become unstable. In online training under TIM, the ordering reverses: Epθ[R∇θlog⁡pθ(y)].\mathbb{E}_{p_\theta}\left[R \nabla_\theta \log p_\theta(y)\right].8 rewards are least stable, and group-centered rewards are comparatively more robust.

Figure 5

Figure 5: Reward sign and online staleness interact differently: offline training favors Epθ[R∇θlog⁡pθ(y)].\mathbb{E}_{p_\theta}\left[R \nabla_\theta \log p_\theta(y)\right].9 rewards, whereas online training under TIM is least stable with qθq_\theta0 rewards.

The implication is that reward centering should not be interpreted as eliminating drift. Group centering makes advantages sum to zero across rollouts for a prompt, but drift is a prefix-level quantity. A prefix can have positive or negative expected advantage even when the group mean is zero, so the conditional expected score remains nonzero. Group centering reduces the magnitude of drift but does not cancel it. SC addresses this residual directly.

Limitations and open questions

The paper’s principal limitation is that SC does not fully correct the off-policy distribution. After drift cancellation, the covariance is still taken under qθq_\theta1 rather than qθq_\theta2. The staleness experiments demonstrate the practical consequence: SC performs better when combined with TIS or MIS. Thus the method’s strongest claim is specifically about removing drift, not about producing an unbiased estimator of the ideal on-policy policy gradient under arbitrary mismatch.

The experimental regime is also deliberately nonstandard. The headline comparisons use severe quantization, fixed sampler-weight noise, or a sampler updated only every 64 steps, often with short sequences. These conditions are intended to approximate the cumulative effect of milder mismatch over longer runs, but the equivalence is not established. The paper leaves open how SC compares under realistic asynchronous systems with continuously batched inference, longer agentic trajectories, mixed sources of mismatch, and smaller per-step numerical errors.

The top-qθq_\theta3 implementation relies on a tail model derived from the trainer distribution. Although qθq_\theta4 and qθq_\theta5 match full SC in all reported settings, the approximation has not been stress-tested across broader distributions or tasks. The largest-scale comparisons also use uneven seed counts, including single-seed results for many baselines, and the MoE experiments do not replay routing decisions. These factors limit the precision with which the reported performance gaps can be attributed solely to the correction rule.

Finally, the paper does not investigate the separate instability of offline negative-reward training in depth. Its explanation in terms of unbounded negative log-probability updates is plausible within the presented setup, but the relationship between that mechanism, SC, reward normalization, and broader off-policy objectives remains unresolved.

Conclusion

The paper identifies a specific failure mode in off-policy LLM RL: trainer-sampler mismatch introduces a nonzero expected score, and the resulting drift distills the trainer toward a biased, moving sampler. Score centering removes this drift through an additive expected-score subtraction that is exact at the prefix level and practical with top-qθq_\theta6 sampler log probabilities.

Across controlled mismatch, severe quantization, sampler staleness, and a 30B MoE model, SC is consistently competitive with importance sampling and substantially more robust in the most severe quantization regimes. Its residual distributional mismatch explains why composition with TIS or MIS is advantageous under staleness. The resulting contribution is both diagnostic and algorithmic: it separates drift cancellation from importance-ratio correction and provides a low-overhead mechanism for stabilizing RL when exact trainer-inference equivalence is impractical.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Explain it Like I'm 14

1. What is this paper about?

This paper studies how to make reinforcement learning (RL) more stable when training LLMs.

In RL, a LLM writes an answer, receives a reward for how good it was, and then changes its behavior to produce better answers next time. The paper focuses on a problem called training-inference mismatch. This happens when the system that creates answers and the system that trains the model are slightly different.

For example, the answer-generating system might use lower-quality number formats, different computer calculations, or an older version of the model. Even tiny differences can eventually make training fail.

The researchers introduce a method called score centering, which is designed to remove the harmful effect of this mismatch.

2. What questions are the researchers asking?

The paper mainly asks:

  • Why can very small differences between the training and answer-generating systems make RL unstable?
  • What exactly causes the model’s performance to get worse over time?
  • Can a simple correction make RL stable without requiring expensive calculations?
  • How does this new correction compare with existing methods, especially importance sampling?
  • Can the correction work when the answer-generating system uses compressed numbers or an older model version?

The researchers’ main idea is that the problem is caused by something they call drift. Drift is a repeated unwanted push that makes the trainer copy the small errors of the answer-generating system.

3. How did the researchers investigate the problem?

Understanding the source of instability

The researchers first studied the mathematical update used in policy-gradient RL. A gradient is a direction that tells the model how to change its settings to improve its rewards—similar to a map showing which way is uphill.

Normally, the model that generates answers and the model that calculates updates are supposed to be identical. However, in real systems:

  • One may use different numerical precision.
  • The two systems may use different computer programs or hardware operations.
  • The answer generator may use an older model checkpoint.
  • The answer generator may be compressed to save memory and run faster.

The researchers separated the training update into two parts:

  1. A useful signal, which shows which choices led to better rewards.
  2. An unwanted drift, which pushes the trainer toward the answer-generating system even when that system’s behavior is not actually better.

An analogy is a student copying from a classmate. If the classmate gives useful answers, copying can help. But if the classmate makes a small mistake, and the student copies that mistake again and again, the mistake can grow.

Creating a correction

The proposed method, score centering, subtracts the average score of all possible next words from the score of the word that was actually chosen.

In simple terms, it asks:

“Was this word unusually helpful compared with the other words the model could have chosen?”

This removes the general push toward the imperfect answer generator and leaves mainly the information about which choices were connected to good or bad rewards.

Comparing different methods

The researchers tested score centering against several existing correction methods based on importance sampling.

Importance sampling is a way to correct for differences between two probability distributions. In this case, it adjusts how much each sampled word counts when the training model and answer-generating model disagree.

However, importance sampling can sometimes give very large weights to rare words, making training noisy and unstable. Practical versions therefore limit or remove extreme weights.

The experiments used:

  • A small model, Qwen3-0.6B, on a task called Countdown.
  • A much larger Qwen3-30B model on mathematics problems from the INTELLECT-2 dataset.
  • Artificial differences between the trainer and sampler.
  • Quantization, which stores numbers in a shorter format to make computation faster.
  • Staleness, where the answer generator is updated only occasionally instead of after every training step.

The researchers also tested whether using only the sampler’s most likely 32 or 128 words was enough for score centering, rather than storing probabilities for the entire vocabulary.

4. What did the researchers find?

Small errors can build up over time

The main finding is that training-inference mismatch creates drift. Each training step may introduce only a tiny unwanted change, but these changes accumulate.

This is like repeatedly pushing a ball slightly to one side. One push may not matter, but after many pushes, the ball can move far from where it should be.

The more severe the mismatch, the sooner the training collapsed.

Score centering removed the harmful drift

Score centering successfully removed this unwanted push in the researchers’ experiments. It allowed the model to focus on the relationship between its choices and their rewards.

This was especially helpful when the sampler used strong quantization or when the sampler was significantly out of date.

It often performed as well as or better than importance sampling

Across the experiments:

  • With mild mismatch, score centering performed about as well as the strongest importance-sampling methods.
  • With severe quantization, score centering was often the only method that remained stable.
  • With severe staleness, combining score centering with importance sampling worked best.
  • On the 30-billion-parameter model, score centering performed especially well under strong compression.

For example, under one very difficult quantization setting, score centering achieved about 30% training accuracy, while truncated importance sampling achieved about 12%, and the other methods finished below 5%.

A small amount of information was enough

The researchers found that they did not need to store the probabilities of every possible next word. Using the top 32 or 128 most likely words gave nearly the same results as using the full vocabulary.

This is important because LLMs may have vocabularies containing more than 100,000 words. Storing every probability for every generated word would require an enormous amount of memory.

5. Why are these findings important?

The paper suggests that RL training does not always need perfectly identical training and inference systems. Making them identical can be expensive and may reduce the speed of training.

Instead, score centering may allow the systems to be somewhat different while keeping training stable. This could make it easier to:

  • Use faster, compressed models for generating answers.
  • Run training and answer generation on separate computers.
  • Use older model versions temporarily.
  • Train LLMs more efficiently.
  • Reduce the amount of expensive hardware engineering needed to keep systems perfectly synchronized.

The method is also relatively simple. It adds a correction to the training calculation rather than using large, unpredictable multiplication factors.

Conclusion

This paper argues that RL training for LLMs can fail because tiny differences between the training model and the answer-generating model create a harmful drift. The trainer slowly copies the sampler’s errors, and those errors can grow into a feedback loop.

The researchers propose score centering, which removes this drift by comparing each chosen word with the average of the possible choices. Their experiments show that this method can make training much more stable, especially when the sampler is heavily compressed or out of date.

The method is not a complete solution for every situation. When the sampler becomes very stale, combining score centering with importance sampling works better. Still, the research could help make future language-model training faster, cheaper, and more reliable.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

The paper leaves the following issues unresolved:

  • Generalization beyond the tested models and tasks: Experiments use primarily Qwen3 models on Countdown and INTELLECT-2 mathematics; it remains unclear whether score centering works similarly for other model architectures, tokenizer vocabularies, domains, reward functions, and environments such as coding, dialogue, tool use, or long-horizon agents.
  • Behavior under realistic, mild mismatch: The headline results rely on deliberately severe quantization, synthetic weight noise, or 64-step sampler staleness. The effectiveness and relative advantage of score centering under the smaller but persistent mismatches typical of production systems are not established.
  • Long-context and long-horizon training: Experiments use maximum sequence lengths of 512 or 1,024 tokens, whereas the paper motivates score centering partly through agentic tasks with episodes lasting minutes, hours, or days. Its stability under long contexts, variable-length trajectories, and highly asynchronous rollouts remains untested.
  • Impact of realistic asynchronous systems: The staleness experiment updates the sampler only every 64 steps, rather than modeling continuous batching, rollout-level delays, subsequence-level staleness, heterogeneous episode durations, or fully disaggregated training and inference.
  • Validity of the drift explanation across RL algorithms: The analysis focuses mainly on REINFORCE with group-centered rewards. It remains unclear how the proposed drift decomposition and score-centering correction behave with PPO, GRPO, DAPO, IcePop, actor–critic methods, outcome-versus-token-level rewards, or generalized advantage estimation.
  • Theoretical convergence guarantees: The paper explains how score centering removes an expected drift term but does not provide convergence, stability, or regret guarantees for the resulting stochastic optimization process, particularly when the sampler is stale, quantized, or updated asynchronously.
  • Consequences of the remaining covariance-distribution mismatch: Score centering estimates the reward–score covariance under the sampler rather than the trainer. The paper observes that this matters under severe staleness but does not characterize the resulting bias, identify conditions under which it is negligible, or quantify how it scales with policy divergence.
  • Optimal composition with importance sampling: Score centering combined with TIS or MIS performs well in some settings, but the paper does not determine when composition is preferable, how to select the importance-weighting function, or whether alternative clipping and masking rules could yield better bias–variance trade-offs.
  • Comparison with fully corrected or jointly engineered systems: The baselines omit components from the original methods, such as dynamic sampling, length penalties, separate PPO surrogates, or complete algorithmic recipes. Consequently, it is unresolved whether score centering outperforms fully configured implementations in end-to-end practice.
  • Sensitivity to optimization choices: Results rely heavily on SGD with a fixed learning rate and one update per batch. The interaction between score centering and Adam/AdamW, learning-rate schedules, multiple epochs, gradient clipping, momentum, batch size, and distributed optimizer behavior is not systematically studied.
  • Dependence on reward and advantage distributions: The analysis emphasizes constant positive rewards and group-centered {−1,+1}\{-1,+1\} rewards. The method’s behavior with sparse, continuous, highly skewed, delayed, noisy, or unbounded rewards—and with imperfect learned critics—remains unknown.
  • Negative-reward and offline-training instability: The paper explicitly leaves offline training with negative rewards unexplored. It does not test whether score centering can mitigate the claimed unbounded-loss or distribution-sharpening mechanisms in offline or mixed offline–online RL.
  • Effect of prompt and group composition: Because group centering is used throughout, the results do not isolate how score centering interacts with group size, heterogeneous prompt difficulty, unequal numbers of valid completions, or alternative advantage-normalization schemes.
  • Top-kk approximation under distributional shift: The top-kk method is validated mainly where the head covers most of the sampler probability mass. Its error and stability when the sampler has a flatter distribution, when sampled tokens frequently fall in the tail, or when trainer and sampler tails differ substantially are not quantified.
  • Principled selection of kk: The paper reports that k=32k=32 and k=128k=128 work in the tested settings but provides no adaptive rule, error bound, or compute–accuracy trade-off for choosing kk across models, sequence positions, temperatures, or mismatch levels.
  • Numerical robustness of tail-mass estimation: The implementation floors tail masses for numerical stability, but the effects of this heuristic under extreme quantization, very peaked distributions, large vocabularies, or near-zero tail mass are not analyzed.
  • Computational and systems cost at production scale: The reported overhead is measured with a custom JAX sampler and does not include implementations in widely used systems such as vLLM or SGLang. Memory, communication, latency, and throughput costs under multi-node serving and large batch sizes remain unresolved.
  • Mixture-of-experts routing effects: The 30B experiment does not replay router indices, despite acknowledging that sampler and trainer routing can differ. The separate contribution of routing mismatch—and whether score centering remains effective when routing decisions are synchronized or corrected—is not isolated.
  • Mismatch sources beyond quantization and staleness: The experiments do not systematically evaluate kernel non-associativity, different attention implementations, activation precision, KV-cache formats, tokenizer discrepancies, independent inference codebases, or implementation bugs.
  • Statistical reliability and reproducibility: The paper notes that many comparisons are difficult to separate under mild mismatch, yet it provides limited information about seed variability, confidence intervals, failure probabilities, and statistical tests across the full set of methods and settings.
  • Long-term training performance: Most experiments emphasize early collapse or a few hundred steps. It remains unclear whether score centering preserves final reward, sample efficiency, and solution quality over substantially longer training runs after apparent stabilization.
  • Relationship between mismatch magnitude and collapse time: The paper qualitatively links stronger mismatch to earlier collapse but does not establish a quantitative scaling law connecting mismatch measures, accumulated drift, training length, model size, sequence length, and collapse probability.
  • Evaluation beyond training accuracy: Results primarily report training accuracy or stability. Generalization to held-out prompts, robustness to distribution shift, calibration, diversity, reward hacking, and downstream task performance are not evaluated.
  • Potential interaction with exploration: Since score centering removes a component that acts like distillation toward the sampler, it may alter exploration dynamics. Whether it suppresses useful exploration or changes mode collapse and policy entropy is not investigated.
  • Applicability when sampler and trainer parameterizations differ: The derivation assumes that the sampler probabilities can be related to the trainer’s token distribution and that the trainer’s scores are available for all relevant tokens. Its validity for distilled samplers, speculative decoding, external proposal policies, or structurally different inference models remains open.

Practical Applications

Immediate Applications

  • Stabilizing RL fine-tuning pipelines for LLMs (software/AI infrastructure). Integrate top-kk score centering into PPO, GRPO, REINFORCE, DAPO, or related LLM-RL trainers to reduce collapse caused by numerical differences between the training and inference engines. The paper’s implementation requires only sampler top-kk log-probabilities and can be expressed as a scalar autograd loss, making it compatible with existing JAX/PyTorch-style training systems. Dependencies: The sampler must expose top-kk log-probabilities and the trainer must be able to compute the corresponding token distributions. The reported results used k=32k=32 or k=128k=128; other architectures and vocabulary distributions may require validation.
  • Using low-precision inference during RL training. Deploy quantized samplers—such as INT8, FP8, or FP4-weight/KV-cache configurations—while retaining higher precision in the trainer. This can reduce inference memory and increase rollout throughput without necessarily causing the reward collapse observed with naïve policy gradients. Relevant sectors: LLM serving, cloud AI infrastructure, specialized inference hardware, and model training. Dependencies: The experiments deliberately used severe quantization and tested particular Qwen3 models. Production systems should tune quantization schemes, monitor reward and gradient statistics, and verify performance across models and tasks.
  • Reducing training–inference engine engineering requirements. Score centering provides a practical alternative to making training and inference numerically identical. Organizations may use different kernels, batching strategies, precisions, or inference runtimes while compensating algorithmically for their persistent bias. This could reduce the need for expensive batch-invariant kernels or fully synchronized numerical implementations. Dependencies: Score centering mitigates systematic drift but does not guarantee equivalence between engines. Severe implementation bugs, incorrect probability logging, or distributional errors outside the approximated top-kk support may remain harmful.
  • Stabilizing asynchronous and disaggregated RL systems. Apply score centering when rollout generation and gradient computation run on separate workers or accelerators. This is particularly useful in systems that prioritize GPU utilization through continuous batching and asynchronous trainer–sampler synchronization. Dependencies: Under large checkpoint staleness, the paper recommends combining score centering with truncated or masked importance sampling, such as TIS or MIS. The method therefore does not eliminate the need to track checkpoint versions and monitor staleness.
  • Improving long-context and agentic coding RL workflows. Coding agents and other long-running agents often produce episodes over widely varying time scales, making stale rollouts unavoidable. Score centering can be added to the training loop to reduce the feedback loop in which a biased, delayed sampler repeatedly distills its own errors into the trainer. Potential products: distributed coding-agent trainers, asynchronous web-agent training platforms, and RL systems for tool-using LLMs. Dependencies: Long-horizon tasks introduce additional problems—credit assignment, reward sparsity, tool failures, and nonstationary environments—that score centering alone does not address.
  • Combining additive correction with importance sampling. Implement a two-stage correction: use TIS or MIS to partially correct the rollout distribution, then center the resulting weighted score. The paper reports that this composition performs especially well under severe staleness and can outperform either correction alone. Dependencies: Importance-ratio estimates require reliable sampler log-probabilities. Clipping or masking thresholds remain design choices and may introduce bias or variance.
  • Improving operational monitoring of RL training. Treat the estimated expected score, sampler–trainer probability divergence, checkpoint age, and reward drift as diagnostic signals. A rising expected score indicates that the sampler and trainer are systematically mismatched and that drift may be accumulating. Potential workflow: add dashboards and automated alerts for drift magnitude, rollout staleness, quantization mode, reward collapse, and correction-method effectiveness. Dependencies: These metrics must be calibrated against the specific model, reward scale, vocabulary, and optimizer; the paper does not establish universal thresholds.
  • Reducing memory and compute costs in RL experimentation. The top-kk approximation avoids storing full vocabulary distributions for every token. The paper reports negligible measured runtime overhead for its implementation and substantial savings relative to full-distribution score centering. Relevant sectors: academic RL research, startup-scale model training, and resource-constrained AI laboratories. Dependencies: The reported overhead was measured in a custom JAX sampler; performance with vLLM, SGLang, PyTorch, or other runtimes requires independent benchmarking.
  • A practical baseline for academic RL experiments. Researchers can use score centering as a controlled baseline when comparing RL algorithms, quantization levels, inference runtimes, or synchronization strategies. This helps separate failures caused by the learning objective from failures caused by training–inference mismatch. Dependencies: The paper’s main comparisons use Qwen3 models, Countdown, and INTELLECT-2 math data, with intentionally severe mismatch. Broader validation across reward models, environments, model families, and random seeds is still needed.
  • Policy and governance testing for efficient AI training. Organizations or regulators evaluating the energy and cost implications of RL-trained models can treat score centering as an enabling technique for lower-precision and higher-utilization training. It may support more efficient experimentation without assuming that every accelerator must run identical numerical workloads. Dependencies: Efficiency gains should be assessed jointly with model quality, safety evaluations, reproducibility, and potential behavioral changes caused by quantization or stale data.

Long-Term Applications

  • Scalable RL training platforms for frontier LLMs. Incorporate score centering into general-purpose distributed RL platforms that automatically select among vanilla score centering, TIS, MIS, or their composition according to measured quantization error and rollout staleness. Such systems could dynamically trade off inference efficiency against correction strength. Dependencies: This requires robust online estimates of distribution mismatch, stable adaptive policies, and validation at substantially longer training horizons and larger batch sizes than those studied.
  • Algorithmically tolerant hardware–software co-design. Future AI accelerators and inference runtimes could be designed around controlled numerical mismatch rather than exact numerical reproducibility. Hardware might expose efficient top-kk probability logging, sampler tail-mass estimates, and checkpoint-age metadata specifically for corrected RL workloads. Dependencies: The benefit depends on whether correction overhead remains small at production scale and whether hardware-induced errors are approximately distributional and recoverable rather than arbitrary.
  • Automated precision and synchronization scheduling. A training controller could select sampler precision, KV-cache precision, synchronization frequency, and importance-sampling strength based on drift measurements. For example, it might use aggressive FP4 inference when drift is low, increase synchronization when staleness rises, and enable TIS plus score centering when the trainer moves far from the sampler. Dependencies: Such closed-loop control requires reliable stability indicators and could introduce its own nonstationarity. It must be tested against delayed rewards and long-horizon training collapse.
  • RL for robotics and embodied agents. The underlying idea—subtracting an estimated expected score to remove mismatch-induced drift—could potentially be adapted to policy-gradient systems in robotics, simulation, and embodied AI when actions are sampled by one numerical stack and gradients are computed by another. Dependencies: The paper’s exact method relies on summing over the language-model vocabulary, which is tractable for discrete tokens. Continuous-control policies would require a different way to estimate the expected score, so direct transfer is not established.
  • Discrete-action applications outside language modeling. Similar corrections may be useful in recommender systems, dialogue policies, game agents, or discrete decision systems where a sampler and trainer use different approximations or stale policies. Dependencies: In conventional RL, exact expectation over a large action space may be difficult or impossible. The LLM setting is unusually favorable because the next-token distribution is explicitly available and can be approximated using top-kk probabilities.
  • More reliable training of tool-using and multi-agent systems. Score centering could become part of RL methods for agents that call APIs, browse, write code, or coordinate with other agents. These systems naturally generate asynchronous and heterogeneous rollouts, making staleness and engine mismatch more severe. Dependencies: Further work is needed to handle tool-induced changes in the environment, partial observability, multi-agent nonstationarity, and reward attribution across very long trajectories.
  • Theoretical development of low-variance off-policy corrections. The paper suggests a broader research direction: additive corrections that remove systematic off-policy drift without multiplying gradients by heavy-tailed importance ratios. Future methods may combine score centering with learned critics, control variates, adaptive tail models, or distributionally robust estimators. Dependencies: Score centering does not fully recover the on-policy objective: its covariance is measured under the sampler distribution rather than the trainer distribution. New methods must address this residual bias while preserving the variance and memory advantages.
  • Certified stability and reproducibility standards for LLM RL. Training systems could eventually report standardized mismatch measures, drift estimates, sampler staleness, quantization configuration, and correction rules alongside model checkpoints. This would make RL results more reproducible across hardware, kernels, and inference engines. Dependencies: The field would need agreed metrics and thresholds. The paper demonstrates the mechanism in selected settings but does not yet provide a formal stability guarantee or a universal reproducibility protocol.
  • Extension to broader reward and policy-optimization objectives. Because the correction is expressed as an additive modification to the score term, it may be integrated into actor–critic, preference optimization, group-based RL, and other policy-gradient objectives. A generalized implementation could center weighted scores after clipping, masking, or other policy corrections. Dependencies: Each objective may have different normalization, advantage, clipping, and sequence-level semantics. The correction must be re-derived and tested rather than assumed to preserve the behavior of every algorithm.
  • Safety-critical deployment of RL-trained models. More stable RL training could reduce accidental reward collapse, but score centering might eventually support safety-oriented training workflows in healthcare, finance, education, and public-sector applications where unstable optimization is costly. Dependencies: Stability is not equivalent to safety. Any such deployment would still require domain-specific validation, robust reward design, adversarial testing, human oversight, and monitoring for reward hacking or distribution shift.

Glossary

  • Advantage: A scalar measuring how much better an action or outcome is than a reference expectation, used to weight policy-gradient updates. “advantages almost always include negative values”
  • Autoregressive sampling: Generating a sequence one token at a time, conditioning each token on previously generated tokens. “Since sampling is autoregressive”
  • Batch-invariant kernel: A computational kernel whose numerical behavior does not depend on the input batch or sequence shape. “Batch-invariant kernels eliminate the dependency on input shapes”
  • Bias: A systematic deviation of an estimator or update from its intended value. “bounding the importance ratio, which inevitably introduces bias”
  • Binary total variation: A distance or discrepancy measure specialized to binary outcomes, used here to define a masking region. “DPPO, which uses a masking region based on binary total variation”
  • Control variate: A correlated, known-mean quantity used to reduce the variance of a stochastic estimator. “reward baselines and score-function control variates”
  • Covariance: A measure of how two random quantities vary together, appearing in the decomposition of the policy-gradient update. “the covariance term sees which token led to which reward”
  • Critic: A learned estimator of expected future reward, commonly used to provide value estimates or baselines in reinforcement learning. “score centering obtains the same expected update without a critic”
  • Cross-entropy loss: A loss measuring the discrepancy between a target probability distribution and a model distribution. “the negative gradient of the cross-entropy (SFT) loss”
  • Distribution sharpening: A process in which probability mass becomes increasingly concentrated on a smaller set of outcomes. “the main mechanism behind this instability is the unboundedness of negative rewards in combination with distribution sharpening”
  • Disaggregating: Separating components of a system so that they can operate independently, often asynchronously. “requires disaggregating training and inference engines”
  • Distillation: Training a student model to reproduce the behavior or output distribution of a teacher model. “vanilla policy gradient distills the trainer toward the sampler”
  • Drift: A systematic, mismatch-induced component of an update that pushes the trainer toward the sampler independently of the task-specific learning signal. “the instability of RL under TIM is primarily caused by drift”
  • Dynamic sampling: Adaptively changing the sampling procedure during training based on current model or training conditions. “DAPO combines asymmetric clipping with dynamic sampling”
  • Entropy: A measure of uncertainty in a probability distribution, although the paper uses the related notion of probability concentration when discussing model behavior. “the tail model of the sampler's distribution”
  • Floating-point non-associativity: The property that changing the order of floating-point operations can change the numerical result. “they might still produce different outputs due to non-associativity of floating point operations”
  • Forward pass: A computation that maps model inputs to outputs without necessarily computing parameter gradients. “requires two separate forward passes through the model”
  • Gradient estimate: A stochastic approximation to the gradient of an objective, often computed from sampled data. “inflating the variance of the gradient estimate”
  • Gradient variance: The variability of stochastic gradient estimates around their expected value. “the reason to center rewards is that it reduces the gradient variance”
  • Importance sampling: A technique that reweights samples from one probability distribution to estimate expectations under another distribution. “we can correct for the training-inference mismatch exactly using importance sampling”
  • Importance ratio: The ratio of a target distribution’s probability to a sampling distribution’s probability, used as a multiplicative correction. “The weighted score gets multiplied by the importance ratio”
  • Inference engine: The system responsible for executing a model to generate sampled outputs, typically optimized for efficient prediction. “one forward pass of the sampler (inference engine)”
  • Isotropic Gaussian distribution: A multivariate Gaussian distribution with equal variance in every direction and no directional covariance. “Δθ\Delta \theta was sampled at the beginning of training from an isotropic Gaussian distribution”
  • Kernel: A low-level computational routine, often optimized for hardware such as GPUs. “even if both engines are correct and use the same GPU kernels”
  • Logprob: The logarithm of the probability assigned to an outcome by a model. “We therefore log only its top-kk logprobs”
  • Maximum likelihood estimator: An estimator that selects parameters maximizing the likelihood of observed data. “SFT with +1/0+1/0 converges to a maximum likelihood estimator”
  • Mixture-of-experts: A neural architecture that routes each input through a selected subset of specialized subnetworks called experts. “Since Qwen3-30B-A3B is a mixture-of-experts model”
  • Off-policy reinforcement learning: Reinforcement learning in which data are generated by a policy different from the policy being optimized. “In contrast, score centering ... uses no importance ratios”
  • On-policy reinforcement learning: Reinforcement learning in which samples are generated by the same policy whose objective is being optimized. “on-policy the expected gradient is zero at every prefix”
  • Policy gradient: A reinforcement-learning method that directly optimizes policy parameters using reward-weighted gradients of log probabilities. “the policy gradient method computes the gradient of the expected reward”
  • Positive feedback loop: A self-reinforcing process in which an initial discrepancy causes subsequent updates that amplify the discrepancy. “the error compounds in a feedback loop”
  • Prefix: The portion of a generated sequence preceding a particular token. “conditional on its prefix”
  • Quantization: Representing model parameters, activations, or cached values with reduced numerical precision. “we test intentionally severe quantization and staleness settings”
  • Rollout: A sequence of actions or tokens sampled by a policy and used as an experience or training example. “RL training of LLMs via policy gradient consists of sampling rollouts”
  • Sampler: The model or inference system that generates sequences used for reinforcement-learning updates. “rollouts are generated from a sampler”
  • Score function: The gradient of the logarithm of a probability with respect to model parameters. “sv=∇θlog⁡pvs_v = \nabla_\theta \log p_v is its score”
  • Score centering: A correction that subtracts the sampler-expected score from each token score to eliminate mismatch-induced drift. “We achieve this simply by subtracting from each score the expected score under the sampler”
  • Score-function estimator: A gradient estimator based on the gradient of log probability, often used for stochastic policies. “score-function control variates”
  • Staleness: The discrepancy caused by using samples generated from an older model checkpoint or policy. “Staleness becomes most severe in long-context environments”
  • Stop-gradient: An automatic-differentiation operation that treats a value as constant and prevents gradients from flowing through it. “sg⁡\operatorname{sg} denotes stop-gradient”
  • Tail mass: The total probability assigned to all outcomes outside a selected high-probability subset. “where ρ\rho is the ratio of the sampler's to the trainer's tail mass”
  • Top-kk approximation: An approximation that retains only the kk highest-probability outcomes and models the remainder collectively or approximately. “we only store a top-kk approximation of the sampler's distribution”
  • Training-inference mismatch (TIM): A discrepancy between the model computations used for training and those used to generate samples. “resulting in the training-inference mismatch (TIM)”
  • Truncated importance sampling: Importance sampling in which ratios exceeding a specified threshold are clipped or otherwise bounded. “score centering can be composed on top of importance sampling methods such as TIS and MIS”
  • Variance reduction: A method for decreasing the variability of stochastic estimates without changing, or while controlling, their expected value. “Subtracting a zero-mean quantity from the policy gradient is a classical variance-reduction idea”
  • Vocabulary distribution: The probability distribution over all tokens that a LLM can generate at a given position. “the expected score under this reconstructed q^\hat{q} distribution”
  • Weight staleness: A mismatch arising because one system uses model parameters that have not yet incorporated recent updates from another system. “the inference engine gets updated only every 64 steps”

Open Problems

We found no open problems mentioned in this paper.

Tweets

Sign up for free to view the 7 tweets with 130 likes about this paper.