Understanding the interaction between CPT and SFT
Characterize the interaction between continued pretraining (CPT) and supervised finetuning (SFT) for long-context vision-language models and determine why these procedures do not compose additively across many benchmarks.
References
The interaction between CPT and SFT remains incompletely understood: they do not compose additively across many benchmarks, suggesting opportunities for mixed-stage training or replay mechanisms.
The relationship between the two stages is therefore largely unexplored, and we believe it is the most promising next step: how the composition of the SFT corpus interacts with the mid-training mixture, whether stronger or larger post-training supervision substitutes for or compounds with a tool-use prior, and how the choice of RL environments and reward design shifts what the prior is worth.