Assess robustness of results under advanced post-training methods
Determine whether the alignment pretraining effects reported for supervised fine-tuning and direct preference optimization persist, diminish, or change when applying reinforcement learning with verifiable rewards, reasoning-focused post-training, deliberative alignment, or constitutional AI.
References
It is unclear whether our findings would be significantly affected by the implementations of these techniques.
Taken together, these results suggest that, in our setup, knowledge-aligned data construction can achieve gains comparable to additional post-training stages; whether these gains stack with factuality-oriented RL remains open.
Two clear open problems remain: determining whether persona binding survives RL and, if not, designing RL recipes that preserve or actively reinforce it.