Characterize the task-dependent reward–language-drift trade-off
Characterize the task-dependent steepness of the relationship between language drift and maximum expected reward during reinforcement learning with verifiable reward, as represented by the curve in Figure 5.
References
In other words, we do not establish the (likely task-dependent) steepness of the curve in Figure 1.
— On Language Drift during RLVR Post-Training
(2610.02015 - Sullivan et al., 1 Oct 2026) in Section 4.3, “Hypothesis: Novel Tasks Induce Language Drift,” p. 5
It is unclear to what degree language mixing represents genuine language drift in the sense of Definition 1, as language mixing is also well-documented in bilingual human language users (in the form of code-switching; see e.g. Poplack, 1980).
— On Language Drift during RLVR Post-Training
(2610.02015 - Sullivan et al., 1 Oct 2026) in Footnote 3, Section 2, “Related Work,” p. 3