Quantify run-to-run variability through multi-seed replication

Quantify the run-to-run variability of the Qwen3.5-2B and Qwen3.5-4B procedural-skill SFT results by replicating the v1.9 and v2.0 training and evaluation experiments across three to five random seeds.

Background

The headline results are based on single-seed training and evaluation, with paired statistical tests performed within that seed. Consequently, the reported procedural-Δ lifts and significance levels do not establish how sensitive the findings are to initialization, data ordering, or other stochastic training effects.

The paper specifically identifies three-to-five-seed replication of the 2B and 4B experiments as the needed test for bounding run-to-run variance.

References

3--5 seed replication of v1.9 and v2.0 would bound run-to-run variance \citep{biderman2024lessons}; we did not run it.

— Procedural-skill SFT across capacity tiers: A W-Shaped pre-SFT Trajectory and Regime-Asymmetric Mechanism on 0.8B-4B Qwen3.5 Models  (2605.11907 - Strozzi, 12 May 2026) in Section 7, “Threats to Validity and Limitations,” item 1