Impact of evaluation awareness on estimating true deployment misalignment

Ascertain whether models’ awareness that they are being evaluated, rather than deployed, causes measured misalignment in evaluation settings to overstate or understate misalignment that would manifest in real deployment contexts, and quantify the direction and magnitude of this effect.

Background

The paper’s evaluations include both chat-like tests and a realistic Claude Code-based sabotage scenario, but the authors note that models may recognize these are not real deployment settings. They consider this a potential disanalogy that could affect the validity of misalignment measurements.

They explicitly state that it is unclear whether evaluation awareness leads to overestimation or underestimation of “true” misalignment and identify this as a remaining source of uncertainty.

References

We have attempted to mitigate this with our code sabotage evaluation, and it isn't clear whether evaluation awareness would cause our results to overstate or understate "true" misalignment.

— Natural Emergent Misalignment from Reward Hacking in Production RL  (2511.18397 - MacDiarmid et al., 23 Nov 2025) in Section 1 (Introduction), Limitations, item 3

We could apply only that correction to their set; the direction control of \cref{sec:floor} would need their contrastive sets and forward passes over 15 models up to 70B, so the stronger check remains open, and the result stands against the one we can run.

— A Probe Direction Is a Property of Its Prompt  (2608.13329 - Noël, 13 Aug 2026) in Section 7, Reanalysis (Section 7; labeled \cref{sec:reanalysis})

The underlying limitation is deeper than benchmark contamination: we do not yet have a reliable one-to-one mapping between the model's internal representation of its situation, what it verbalizes about that situation, and how it subsequently acts.

— Xeno-Interpretability: Investigating the Alien Minds of LLMs  (2609.20408 - Pierucci et al., 17 Sep 2026) in Section 7, subsection “Evaluation awareness as a measurement problem”

More fundamentally, measuring evaluation awareness is an unsolved problem. We cannot measure unverbalized evaluation awareness directly, and instead use the realism win rate as a metric of realism to complement the measure of verbalized awareness that likely underestimates the models' evaluation awareness. However, gaining high confidence that target models are not recognizing they are being evaluated and acting on that remains methodologically open.

— Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds  (2609.02302 - Ahlqvist et al., 2 Sep 2026) in Section “Limitations and Conclusion,” subsection “Limitations”

We cannot rule out that VEA reduces the measured amount of shutdown sabotage, and hence the amount of misalignment we observe (see Figure~\ref{fig:sabotage-x-eval-awareness}).

— Shutdown Sabotage Propensities in Multi-Agent Systems  (2609.28274 - Knecht et al., 23 Sep 2026) in Section Discussion, limitations (3)