Impact of evaluation awareness on estimating true deployment misalignment
Ascertain whether models’ awareness that they are being evaluated, rather than deployed, causes measured misalignment in evaluation settings to overstate or understate misalignment that would manifest in real deployment contexts, and quantify the direction and magnitude of this effect.
References
We have attempted to mitigate this with our code sabotage evaluation, and it isn't clear whether evaluation awareness would cause our results to overstate or understate "true" misalignment.
We could apply only that correction to their set; the direction control of \cref{sec:floor} would need their contrastive sets and forward passes over 15 models up to 70B, so the stronger check remains open, and the result stands against the one we can run.
The underlying limitation is deeper than benchmark contamination: we do not yet have a reliable one-to-one mapping between the model's internal representation of its situation, what it verbalizes about that situation, and how it subsequently acts.
More fundamentally, measuring evaluation awareness is an unsolved problem. We cannot measure unverbalized evaluation awareness directly, and instead use the realism win rate as a metric of realism to complement the measure of verbalized awareness that likely underestimates the models' evaluation awareness. However, gaining high confidence that target models are not recognizing they are being evaluated and acting on that remains methodologically open.
We cannot rule out that VEA reduces the measured amount of shutdown sabotage, and hence the amount of misalignment we observe (see Figure~\ref{fig:sabotage-x-eval-awareness}).