Effect of alternative OPD objectives on pass@k coverage and RL interaction
Investigate whether alternative on-policy distillation objectives—including reverse KL over the full next-token distribution or a top-$K$ truncation, forward KL, and Jensen–Shannon divergence—expand pass@$k$ coverage more effectively than the token-level reverse-KL objective and determine how the resulting cold starts interact with subsequent reinforcement learning.
References
A natural direction for future work is to investigate alternative OPD objectives that may yield a different cold-start for RL, including reverse KL evaluated over the full next-token distribution (or top-$K$), as well as forward KL or Jensen-Shannon divergence (JSD) objectives, which differ from reverse KL in their mode-coverage behavior and effect on generation diversity~\citep{eopd}. Whether these objectives expand pass@$k$ coverage more effectively, and how the resulting cold-start interacts with RL, remain open questions.