Establish the specificity of procedural-skill SFT relative to generic instruction tuning
Establish whether fine-tuning the same Qwen3.5 base models on 353 rows of generic Opus instruction-tuning data produces the same, smaller, negative, or otherwise different cross-scale procedural-Δ lift as procedural-skill SFT.
References
The specificity claim --- that procedural-skill SFT delivers the $+0.04$ to $+0.075$ Î-lift, as opposed to any 353-row Opus-trace SFT --- is unwitnessed by the most natural ablation: training the same Qwen3.5 base on 353 rows of generic Opus instruction-tuning data and re-running the bench.
— Procedural-skill SFT across capacity tiers: A W-Shaped pre-SFT Trajectory and Regime-Asymmetric Mechanism on 0.8B-4B Qwen3.5 Models
(2605.11907 - Strozzi, 12 May 2026) in Section 6, subsection “What we did not separately decompose”; Section 7, “Threats to Validity and Limitations,” item 5