Derive scaling laws for alignment pretraining
Determine the precise scaling behavior of alignment pretraining interventions as a function of model size, data quantity, and training compute, including whether small fixed data mixtures can reliably influence alignment priors at scale and how effects interact with increased post-training FLOPS.
References
Although evidence from suggests that the effect of pretraining priors increases with model size and data quantity, the precise scaling behaviour of safety interventions at pretraining remains unknown.
Our experiments are conducted at relatively small scales compared with frontier models, and whether our findings hold at frontier scale remains an open question.
For the risks that matter most for frontier assistants, which world we are in remains an open question until the protocol of \cref{sec:protocol}, or something like it, is carried out at their scale.