Characterize the effect of teacher-model scale
Determine how the scale of the teacher model interacts with the performance of Update-aware Sharpness-Aware Minimization (USA) for cross-domain on-policy distillation and model merging.
References
The teachers of our main experiments are moreover all $14$B models and we vary the student instead, since holding the teacher scale fixed keeps each result differing from the others in the student alone, and how the teacher scale interacts with the effect reported here is likewise left to future work.
— USA: Update-aware SAM for Cross-domain On-Policy Disitllation of Language Agents
(2609.34225 - Zhong et al., 28 Sep 2026) in Appendix, Section 1, Limitations (Appendix A)