Characterize the effect of teacher-model scale

Determine how the scale of the teacher model interacts with the performance of Update-aware Sharpness-Aware Minimization (USA) for cross-domain on-policy distillation and model merging.

Background

The main experiments fix the teacher models at 14 billion parameters while varying the student scale, allowing the study to isolate student-side effects. Consequently, the interaction between teacher size and USA's effectiveness remains unexamined. The paper explicitly leaves this interaction for future work.

References

The teachers of our main experiments are moreover all $14$B models and we vary the student instead, since holding the teacher scale fixed keeps each result differing from the others in the student alone, and how the teacher scale interacts with the effect reported here is likewise left to future work.

— USA: Update-aware SAM for Cross-domain On-Policy Disitllation of Language Agents  (2609.34225 - Zhong et al., 28 Sep 2026) in Appendix, Section 1, Limitations (Appendix A)