Establish the specificity of procedural-skill SFT relative to generic instruction tuning

Establish whether fine-tuning the same Qwen3.5 base models on 353 rows of generic Opus instruction-tuning data produces the same, smaller, negative, or otherwise different cross-scale procedural-Δ lift as procedural-skill SFT.

Background

The paper attributes a +0.040 to +0.075 procedural-Δ lift to SFT on demonstrations containing explicit procedural skills. However, the experiments do not include a matched control trained on generic instruction-tuning traces, so the observed improvement could reflect general chain-of-thought or response-discipline learning rather than procedural specificity.

The missing ablation is concrete: train the same Qwen3.5 base at the relevant scales on 353 generic Opus instruction-tuning examples and rerun the procedural bench. The paper identifies several unresolved possible outcomes, including an equal lift, a smaller or negative lift, or a different cross-scale trajectory.

References

The specificity claim --- that procedural-skill SFT delivers the $+0.04$ to $+0.075$ Δ-lift, as opposed to any 353-row Opus-trace SFT --- is unwitnessed by the most natural ablation: training the same Qwen3.5 base on 353 rows of generic Opus instruction-tuning data and re-running the bench.

— Procedural-skill SFT across capacity tiers: A W-Shaped pre-SFT Trajectory and Regime-Asymmetric Mechanism on 0.8B-4B Qwen3.5 Models  (2605.11907 - Strozzi, 12 May 2026) in Section 6, subsection “What we did not separately decompose”; Section 7, “Threats to Validity and Limitations,” item 5