Quantify run-to-run variability through multi-seed replication
Quantify the run-to-run variability of the Qwen3.5-2B and Qwen3.5-4B procedural-skill SFT results by replicating the v1.9 and v2.0 training and evaluation experiments across three to five random seeds.
References
3--5 seed replication of v1.9 and v2.0 would bound run-to-run variance \citep{biderman2024lessons}; we did not run it.
— Procedural-skill SFT across capacity tiers: A W-Shaped pre-SFT Trajectory and Regime-Asymmetric Mechanism on 0.8B-4B Qwen3.5 Models
(2605.11907 - Strozzi, 12 May 2026) in Section 7, “Threats to Validity and Limitations,” item 1