Papers
Topics
Authors
Recent
Search
2000 character limit reached

T2S: A Rehearsal-Based Approach for Extraction-Resistant Model Watermarking

Published 10 Jun 2026 in cs.CR and cs.AI | (2606.11698v1)

Abstract: Model watermarking safeguards AI model intellectual property by embedding distinctive knowledge that induces unique behavioral signatures. The primary technical challenge lies in ensuring watermark robustness against various post-processing attacks on the watermarked model. Model extraction attacks emerge as the most severe threat, where adversaries exploit prediction outputs to train surrogate models that illegally replicate the original model's functionality. In this work, we propose a rehearsal-based watermark embedding framework to enhance the robustness of model watermarks against model extraction attacks. By simulating the extraction process, our method leverages the loss of a \textit{simulated stolen model} on a trigger set as a training signal to fine-tune the watermark knowledge within the target model. This fine-tuning step encourages the watermark to be embedded in a way that boosts transferability, thereby increasing its chances of persisting and remaining detectable in stolen models. Comprehensive experiments conducted under diverse settings demonstrate that the proposed method significantly improves the robustness of model watermarks against both model extraction and subsequent watermark removal attacks.

Summary

  • The paper introduces T2S, a rehearsal-based bi-level optimization method that fine-tunes a target model using feedback from a simulated stolen model, achieving up to 97.56% watermark success under hard-label Knockoff extraction on CIFAR-10.
  • T2S preserves watermarks across Knockoff and DFME attacks, second-round extraction, quantization, and pruning while limiting target-model accuracy loss to approximately one percentage point.
  • The method works with OOD, blended-image, and feature-based triggers without independently trained negative models, but its second-order gradients create substantial memory costs and limit demonstrated scalability beyond moderate-sized vision models.

Motivation and problem statement

Black-box model extraction attacks, in which an adversary queries a deployed classifier and trains a surrogate on the returned predictions, remain the most severe threat to the intellectual property of deep neural networks. Passive defenses based on black-box watermarking embed trigger-set knowledge so that ownership can be verified by querying a suspect model. The central difficulty is that watermarks embedded via conventional training are largely destroyed during extraction: if watermark features decouple from those supporting the classification task, the neurons activated by triggers do not transfer to the stolen model (2606.11698). Prior mitigation strategies either entangle trigger samples with training data in feature space (EWE) or optimize the trigger set itself using a simulated stolen model (SSW). The latter, however, relays feedback only indirectly through an evolving trigger set and requires independently pre-trained negative models to preserve watermark uniqueness.

The T2S framework

T2S ("Tuning to Survive") replaces SSW's indirect pathway with direct meta-optimization of the target model using feedback from a simulated stolen model (SSM). The pipeline has two stages:

  1. Pre-training: a watermark-free model θC\theta_C is trained on D\mathcal{D}; it initializes the SSM (θS←θC\theta_S \leftarrow \theta_C). Separately, a vanilla watermarked target θT\theta_T is trained on D∪T\mathcal{D} \cup \mathcal{T} with cross-entropy over mixed batches.
  2. Rehearsal-based fine-tuning: alternating updates in which the SSM distills from θT\theta_T via KL divergence on training batches, and the target model is then updated with the gradient of the SSM's watermark loss LWM\mathcal{L}_{WM} on the fixed trigger set, combined with a utility loss weighted by coefficient α\alpha.

Because θS\theta_S is a function of θT\theta_T through the distillation update, the chain rule yields second-order derivatives of D\mathcal{D}0 with respect to D\mathcal{D}1 — a bi-level optimization analogous to meta-learning formulations such as MAML-style unrolling. Notably, since the trigger set is fixed, no independently trained negative models are needed to guarantee low false positives, unlike SSW.

The paper demonstrates that this phased design is necessary: training both models from scratch fails to converge (CIFAR-10 plateaus at 36.47% ACC and 23.00% WSR after 30 epochs), whereas fine-tuning raises the actual stolen model's WSR sharply within one epoch while clean accuracy recovers subsequently. The framework is trigger-agnostic: experiments instantiate it with OOD, Mix (blended-image), and feature-based (Narcissus-style optimized) trigger sets.

Robustness against model extraction

The headline results concern watermark retention under Knockoff (soft- and hard-label) and DFME extraction from ResNet-18 targets on CIFAR-10, CIFAR-100, and Tiny-ImageNet. Representative WSRs of stolen models under Knockoff hard-label — the hardest setting evaluated:

Method CIFAR-10 CIFAR-100 Tiny-ImageNet
Content 0.00 0.00 0.00
EWE 0.60 3.62 10.37
SSW-S 72.40 34.75 88.02
T2S 97.56 68.68 90.40

Under soft-label Knockoff, T2S reaches 99.87% WSR on CIFAR-10 with only 0.23 standard deviation, versus 89.60% for SSW-S with substantially higher variance. All methods achieve near-100% WSR against DFME, indicating that data-free extraction is comparatively easy to survive. Watermarking costs at most ~1 point of target-model accuracy across all datasets, and all methods achieve 100% WSR on the target model itself — confirming that embedding, not retention, is the easy part.

Two robustness analyses strengthen these claims. Cross-architecture extraction (ResNet-18 targets stolen into WRN-16-4, ViT-s, MobileViT-s) preserves high WSRs for convolutional surrogates, though transformer surrogates yield lower absolute accuracy due to data-hungry training rather than watermark loss. Varying the query dataset shows that out-of-distribution queries (STL-10, Tiny-ImageNet) actually produce higher WSRs than in-distribution ones; the authors attribute this to the optimization trade-off between primary-task learning and watermark encoding being absent when queries come from unseen distributions.

Persistence under removal attacks

The transferred watermark survives post-theft removal attempts. Second-round Knockoff extraction changes ACC and WSR by less than 0.5%. On stolen models, WSR stays above 95% under 4-bit quantization and above 90% until pruning exceeds 70%, at which point clean accuracy also collapses — meaning the watermark degrades no faster than utility. Against compression of the target model before deployment, T2S retains 100% WSR even at 4-bit quantization. Distillation-based compression is the weakest point: ground-truth labels in the student's training compromise all methods' watermarks, but T2S still leads SSW-S by over 13 points on CIFAR-10 (73.87% vs. 60.00%) and over 35 points on CIFAR-100 (80.67% vs. 45.40%).

Ablations confirm that the rehearsal mechanism, not the trigger design, carries the robustness: without fine-tuning, WSR falls to 3.73–52.67% depending on trigger type, while the default setting yields 94.47–99.87% across all three trigger types. Feature-based triggers provide a consistent but secondary advantage. An exhaustive sweep of source–target class pairs for feature-based triggers on CIFAR-10 yields average stolen-model WSR of 99.27 ± 2.51, with the worst pair ("cat"→"dog") still reaching 78%.

Limitations and open questions

The paper is candid about computational cost. The second-order derivative computation dominates memory: fine-tuning ResNet-18 on CIFAR-10 requires 425 seconds and 4.9 GB per epoch at batch size 100, rising to 17.6 GB at batch size 1000 on a single RTX 3090. The authors position the method as suited to edge-deployed models of moderate scale and suggest gradient checkpointing or multi-GPU training for larger models, but do not demonstrate scaling beyond ResNet-18-class architectures. Two further caveats bear on interpretation. First, the simulation assumes a worst-case attacker who knows the architecture and holds the training data; behavior under weaker or stronger adversaries (e.g., adaptive attacks aware of the rehearsal mechanism) is not evaluated. Second, comparison with MEA-Defender was abandoned because the authors could not reproduce its reported results reliably — their reproduction yielded lower WSR than any value reported in the original paper — leaving that baseline unresolved. Whether the second-order fine-tuning remains tractable and effective for large vision transformers or foundation-scale models remains an open question.

Conclusion

T2S reformulates simulation-based watermarking as direct bi-level fine-tuning of the target model against a simulated stolen model, eliminating the need for evolving trigger sets and independently trained negative models. It delivers near-perfect watermark retention under Knockoff extraction (97.56% WSR hard-label on CIFAR-10), survives second extraction, quantization, pruning, and distillation better than prior rehearsal-based baselines, and does so with negligible accuracy cost and without false-positive-inducing auxiliary models. Its principal constraints are the GPU memory demands of second-order gradients and evaluation confined to moderate-scale image classifiers.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.