Assessing Neural Network Robustness via Adversarial Pivotal Tuning
Abstract: The robustness of image classifiers is essential to their deployment in the real world. The ability to assess this resilience to manipulations or deviations from the training data is thus crucial. These modifications have traditionally consisted of minimal changes that still manage to fool classifiers, and modern approaches are increasingly robust to them. Semantic manipulations that modify elements of an image in meaningful ways have thus gained traction for this purpose. However, they have primarily been limited to style, color, or attribute changes. While expressive, these manipulations do not make use of the full capabilities of a pretrained generative model. In this work, we aim to bridge this gap. We show how a pretrained image generator can be used to semantically manipulate images in a detailed, diverse, and photorealistic way while still preserving the class of the original image. Inspired by recent GAN-based image inversion methods, we propose a method called Adversarial Pivotal Tuning (APT). Given an image, APT first finds a pivot latent space input that reconstructs the image using a pretrained generator. It then adjusts the generator's weights to create small yet semantic manipulations in order to fool a pretrained classifier. APT preserves the full expressive editing capabilities of the generative model. We demonstrate that APT is capable of a wide range of class-preserving semantic image manipulations that fool a variety of pretrained classifiers. Finally, we show that classifiers that are robust to other benchmarks are not robust to APT manipulations and suggest a method to improve them. Code available at: https://captaine.github.io/apt/
- Image2stylegan: How to embed images into the stylegan latent space? In Proceedings of the IEEE international conference on computer vision, pages 4432–4441, 2019.
- Image2stylegan++: How to edit the embedded images? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8296–8305, 2020.
- Advances in adversarial attacks and defenses in computer vision: A survey. IEEE Access, 9:155161–155196, 2021.
- Adef: An iterative algorithm to construct adversarial deformations. arXiv preprint arXiv:1804.07729, 2018.
- Strike (with) a pose: Neural networks are easily fooled by strange poses of familiar objects. arXiv preprint arXiv:1811.11553, 2018.
- data2vec: A general framework for self-supervised learning in speech, vision and language. arXiv preprint arXiv:2202.03555, 2022.
- Unrestricted adversarial examples via semantic manipulation. arXiv preprint arXiv:1904.06347, 2019.
- Adversarial patch. arXiv preprint arXiv:1712.09665, 2017.
- Inverting the generator of a generative adversarial network. IEEE transactions on neural networks and learning systems, 30(7):1967–1974, 2018.
- Evaluating robustness to context-sensitive feature perturbations of different granularities. arXiv preprint arXiv:2001.11055, 2020.
- A rotation and a translation suffice: Fooling cnns with simple transformations. arXiv preprint arXiv:1712.02779, 2017.
- Roger Fletcher. Practical methods of optimization. John Wiley & Sons, 2013.
- Achieving robustness in the wild via adversarial mixing with disentangled representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1211–1220, 2020.
- Vision models are more robust and fair when pretrained on uncurated images without supervision, 2022.
- Collaborative learning for faster stylegan embedding. arXiv preprint arXiv:2007.01758, 2020.
- Masked autoencoders are scalable vision learners. arXiv:2111.06377, 2021.
- Deep residual learning for image recognition. CoRR, abs/1512.03385, 2015.
- Natural adversarial examples. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15262–15271, 2021.
- Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, page 6629–6640, Red Hook, NY, USA, 2017. Curran Associates Inc.
- Semantic adversarial examples. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 1614–1619, 2018.
- Semantic adversarial attacks: Parametric transformations that fool deep classifiers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4773–4783, 2019.
- Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8110–8119, 2020.
- Perceptual adversarial robustness: Defense against unseen threat models. arXiv preprint arXiv:2006.12655, 2020.
- Autoregressive image generation using residual quantization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11523–11532, 2022.
- Dual manifold adversarial robustness: Defense against lp and non-lp adversarial attacks. Advances in Neural Information Processing Systems, 33:3487–3498, 2020.
- Precise recovery of latent vectors from generative adversarial networks. arXiv preprint arXiv:1702.04782, 2017.
- Frequency-driven imperceptible adversarial attack on semantic similarity. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15315–15324, 2022.
- Learning inverse mapping by autoencoder based generative adversarial nets. In International Conference on Neural Information Processing, pages 207–216. Springer, 2017.
- Towards deep learning models resistant to adversarial attacks. arXiv preprint arXiv:1706.06083, 2017.
- Prime: A few primitives can boost robustness to common corruptions. arXiv preprint arXiv:2112.13547, 2021.
- Null-text inversion for editing real images using guided diffusion models, 2022.
- Invertible conditional gans for image editing. arXiv preprint arXiv:1611.06355, 2016.
- Robustness and generalization via generative adversarial training. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15711–15720, 2021.
- Semanticadv: Generating adversarial examples via attribute-conditioned image editing. In European Conference on Computer Vision, pages 19–37. Springer, 2020.
- Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
- Pivotal tuning for latent-based editing of real images. ACM Trans. Graph., 2021.
- High-resolution image synthesis with latent diffusion models, 2021.
- Projected gans converge faster. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
- Stylegan-xl: Scaling stylegan to large diverse datasets, 2022.
- Colorfool: Semantic adversarial colorization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1151–1160, 2020.
- Constructing unrestricted adversarial examples with generative models. In Advances in Neural Information Processing Systems, pages 8312–8323, 2018.
- Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199, 2013.
- Deepfakes and beyond: A survey of face manipulation and fake detection. Information Fusion, 64:131–148, 2020.
- Spatially transformed adversarial examples. In International Conference on Learning Representations, 2018.
- Towards feature space adversarial attack. arXiv preprint arXiv:2004.12385, 2020.
- The unreasonable effectiveness of deep features as a perceptual metric, 2018.
- Understanding the robustness in vision transformers. In International Conference on Machine Learning (ICML), 2022.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.