Optimal teacher-scoring temperature for OPD

Establish whether a teacher-scoring temperature of 0.75 is optimal for on-policy distillation, rather than merely a working choice supported by the reported pilot experiments.

Background

The appendix reports pilot experiments over several teacher-logit scoring temperatures and adopts 0.75 for the main DeepScaleR setting. However, the experiments involve single runs, unequal coverage at one temperature, and limited standalone-teacher sampling, so the reported evidence only motivates a working value and does not resolve the optimization of this hyperparameter.

References

Together with the strongest observed training point at $T=0.75$, these results motivated adopting $0.75$ as the working temperature, but do not establish that it is optimal.

— Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective  (2610.03185 - Cui et al., 2 Oct 2026) in Appendix, Section "DeepScaleR Teacher-Temperature Pilot"