Determinants of the OPD absorption rate

Determine what sets the absorption rate in on-policy distillation, namely the proportion of the remaining teacher–student gap that a single optimization update absorbs.

Background

The paper characterizes on-policy distillation as algorithm-starved: although a single query can expose a broad range of supervised states, the student increasingly slowly absorbs the remaining teacher signal. Experiments show that the absorption rate declines over training and does so similarly across training-set sizes, but the mechanism governing this decline is not identified.

The unresolved issue concerns the factors that determine the absorption rate itself, including why alignment slows even when the training states are fixed. Resolving it could guide the development of more step-efficient OPD optimizers and update rules.

References

What sets the absorption rate is still open.

Rethinking On-Policy Distillation of Large Language Models II: One Training Example  (2609.04172 - Fu et al., 3 Sep 2026) in Section 9, Conclusion, Limitations