Characterize adversarial model modifications
Characterize how an adversary will modify a watermarked text-to-image model after obtaining it, in order to develop a model of watermark persistence that does not rely solely on a bounded parameter-change assumption.
References
Because we do not know exactly how the adversary will modify the model, we denote the model-weight change by $\Delta\beta$ and assume that it is bounded as $|\Delta\beta|_2\leq\rho$.
— Persistent Watermarking of Text-to-Image Models
(2609.39024 - Yao et al., 30 Sep 2026) in Appendix, Section “Benefits of Our Approach on a Toy Problem”