Determine whether self-training causes safety erosion

Determine whether the audited self-training methods cause a practically meaningful change in refusal behavior, using a sufficiently powered and appropriately replicated refusal probe rather than the underpowered ten- or fifty-prompt measurements used in the study.

Background

The refusal probe produces materially different results across repeated evaluations of the same stored checkpoints: one replicate shows a statistically significant decrease in refusal, while neighboring replicates show no significant change. The authors attribute this instability to measurement noise and report only a bound of roughly eleven percentage points.

The paper therefore leaves unresolved whether self-training produces no safety erosion or an effect of practical size. It states that resolving the latter requires a larger probe in which prompts, rather than repeated samples of the same prompts, are increased.

References

We therefore decline to claim that self-training causes no safety erosion here, and equally decline to claim the opposite: what the data support is a bound, that any change is smaller than roughly eleven points, which is compatible both with nothing happening and with an effect of practical size. Detecting the latter would need a probe several times larger, and prompts, not samples, are the unit that has to grow.

Phantom Gains: Auditing Self-Improvement Against a Measured Null  (2608.20290 - Xu et al., 20 Aug 2026) in Section 1, “Introduction,” and Appendix, Section “Out-of-domain probes,” paragraph “Replication settles the point empirically”