Characterize adversarial model modifications

Characterize how an adversary will modify a watermarked text-to-image model after obtaining it, in order to develop a model of watermark persistence that does not rely solely on a bounded parameter-change assumption.

Background

The paper analyzes watermark persistence under adversarial adaptation by representing the adversary’s model-weight change as a perturbation Δβ\Delta\beta bounded by ∥Δβ∥2≤ρ\|\Delta\beta\|_2\leq\rho. The theoretical comparison between the standard trigger-fitting objective and the proposed contrastive-style objective is therefore conditional on this abstract perturbation model.

The authors explicitly acknowledge that the actual form of adversarial modification is unknown. A more complete treatment would characterize the possible transformations used against a watermarked text-to-image model and relate them to the theoretical robustness guarantee, including the extent to which different attacks align with the most damaging perturbation directions.

References

Because we do not know exactly how the adversary will modify the model, we denote the model-weight change by $\Delta\beta$ and assume that it is bounded as $|\Delta\beta|_2\leq\rho$.

— Persistent Watermarking of Text-to-Image Models  (2609.39024 - Yao et al., 30 Sep 2026) in Appendix, Section “Benefits of Our Approach on a Toy Problem”