Make critique refinement neutral to seed adherence

Develop a critique refinement protocol for Petri-style alignment audits that preserves seed adherence across target models while retaining the protocol’s realism improvements.

Background

Critique refinement improves the realism of simulated alignment audits but can cause the auditor to drift away from the seed instruction that defines the intended probe. The magnitude of this degradation varies substantially across target models: it is modest for Sonnet 4.6, Opus 4.8, and Gemini 3.5 Flash but large for GPT-5.5. Although realism gains persist on subsets of seeds with approximately preserved adherence, the paper does not provide a method that eliminates this trade-off.

References

We did not find a way to run the protocol without some loss of seed adherence, and the size of the loss varies by target: it is modest on Sonnet 4.6, Opus 4.8, and Gemini 3.5 Flash, but large on GPT-5.5, where seed adherence drops from 0.92 at baseline to 0.72 at $cr2bo4$ (Fig.~\ref{fig:cost-scaling-grid-pairwise}). Realism gains persist on the subset of seeds where adherence is approximately preserved (at most 0.5 points below baseline; Fig.~\ref{fig:method-comparison-controlled}), but making the protocol adherence-neutral across targets remains open.

Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds  (2609.02302 - Ahlqvist et al., 2 Sep 2026) in Section “Limitations and Conclusion,” subsection “Limitations”