Determine Whether OPSA Scales to Larger Models and Mixture-of-Experts Architectures

Determine whether On-Policy Self-Adaptation (OPSA) scales effectively to models larger than 9 billion parameters and to mixture-of-experts architectures.

Background

The experiments evaluate OPSA only on relatively small models, with the largest model containing 9 billion parameters. Consequently, the evidence does not establish whether OPSA remains effective as model size increases or when applied to mixture-of-experts architectures. The authors explicitly identify this scalability question as unresolved and note that computational limitations prevented them from testing it.

References

It therefore remains unclear whether OPSA scales effectively to larger models or mixture-of-experts architectures.

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement  (2608.31046 - Ding et al., 31 Aug 2026) in Section 6, “Limitations and Future Directions”