Empirical validation of theoretical results in realistic scenarios
Verify which parts of the theoretical convergence and off-policyness results for reward-guided self-training with the RE algorithm hold in realistic scenarios.
References
In terms of empirical work, it remains open to verify which parts of our results hold true in realistic scenarios, see if our theoretical results can inspire better practice of reward-guided self-training, and identify gaps between theory and practice that require further research.
— Fine-Tuning on Self-Generated and Reward-Weighted Data: Learning Dynamics, Convergence Rates, and Benefits of Off-Policyness
(2609.36945 - Wang et al., 29 Sep 2026) in Section 6, “Limitations and future work”