Effect of alternative OPD objectives on pass@k coverage and RL interaction

Investigate whether alternative on-policy distillation objectives—including reverse KL over the full next-token distribution or a top-$K$ truncation, forward KL, and Jensen–Shannon divergence—expand pass@$k$ coverage more effectively than the token-level reverse-KL objective and determine how the resulting cold starts interact with subsequent reinforcement learning.

Background

The OPD stage in the study uses an on-policy reverse-KL objective estimated token by token on student-generated rollouts. The authors point out that other divergence choices can differ in mode coverage and generation diversity, potentially producing different initializations for the later RL stage.

The unresolved questions concern both the capability-expansion effect of these alternatives—specifically whether they improve pass@kk coverage more effectively—and the downstream consequences for RL when each alternative supplies the cold start. The paper explicitly characterizes these as open questions.

References

A natural direction for future work is to investigate alternative OPD objectives that may yield a different cold-start for RL, including reverse KL evaluated over the full next-token distribution (or top-$K$), as well as forward KL or Jensen-Shannon divergence (JSD) objectives, which differ from reverse KL in their mode-coverage behavior and effect on generation diversity~\citep{eopd}. Whether these objectives expand pass@$k$ coverage more effectively, and how the resulting cold-start interacts with RL, remain open questions.

Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR  (2609.04108 - Li et al., 3 Sep 2026) in Limitations, paragraph “Choice of OPD objective”