Characterize the task-dependent reward–language-drift trade-off

Characterize the task-dependent steepness of the relationship between language drift and maximum expected reward during reinforcement learning with verifiable reward, as represented by the curve in Figure 5.

Background

The paper proves that RLVR permits unbounded language drift and that constraining language drift necessarily limits the maximum attainable expected reward. However, these theorems do not determine how much language drift is practically required to achieve a given reward on a particular task and model. The authors explicitly leave unresolved the likely task-dependent steepness of this reward–drift relationship.

References

In other words, we do not establish the (likely task-dependent) steepness of the curve in Figure 1.

— On Language Drift during RLVR Post-Training  (2610.02015 - Sullivan et al., 1 Oct 2026) in Section 4.3, “Hypothesis: Novel Tasks Induce Language Drift,” p. 5

It is unclear to what degree language mixing represents genuine language drift in the sense of Definition 1, as language mixing is also well-documented in bilingual human language users (in the form of code-switching; see e.g. Poplack, 1980).

— On Language Drift during RLVR Post-Training  (2610.02015 - Sullivan et al., 1 Oct 2026) in Footnote 3, Section 2, “Related Work,” p. 3