Dynamic Texture Transfer using PatchMatch and Transformers
Abstract: How to automatically transfer the dynamic texture of a given video to the target still image is a challenging and ongoing problem. In this paper, we propose to handle this task via a simple yet effective model that utilizes both PatchMatch and Transformers. The key idea is to decompose the task of dynamic texture transfer into two stages, where the start frame of the target video with the desired dynamic texture is synthesized in the first stage via a distance map guided texture transfer module based on the PatchMatch algorithm. Then, in the second stage, the synthesized image is decomposed into structure-agnostic patches, according to which their corresponding subsequent patches can be predicted by exploiting the powerful capability of Transformers equipped with VQ-VAE for processing long discrete sequences. After getting all those patches, we apply a Gaussian weighted average merging strategy to smoothly assemble them into each frame of the target stylized video. Experimental results demonstrate the effectiveness and superiority of the proposed method in dynamic texture transfer compared to the state of the art.
- S. Yang, J. Liu, Z. Lian, and Z. Guo, “Awesome typography: Statistics-based text effects transfer,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 7464–7473.
- Y. Men, Z. Lian, Y. Tang, and J. Xiao, “A common framework for interactive texture transfer,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 6353–6362.
- S. Yang, J. Liu, W. Yang, and Z. Guo, “Context-aware text-based binary image stylization and synthesis,” IEEE Transactions on Image Processing, vol. 28, no. 2, pp. 952–964, 2018.
- T. R. Shaham, T. Dekel, and T. Michaeli, “Singan: Learning a generative model from a single natural image,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 4570–4580.
- J. Xie, S.-C. Zhu, and Y. Nian Wu, “Synthesizing dynamic patterns by spatial-temporal generative convnet,” in Proceedings of the ieee conference on computer vision and pattern recognition, 2017, pp. 7093–7101.
- M. Tesfaldet, M. A. Brubaker, and K. G. Derpanis, “Two-stream convolutional networks for dynamic texture synthesis,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 6703–6712.
- A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe, “First order motion model for image animation,” Advances in Neural Information Processing Systems, vol. 32, pp. 7137–7147, 2019.
- C. Thomas, Y. Song, and A. Kovashka, “Learning to transfer visual effects from videos to images,” arXiv preprint arXiv:2012.01642, 2020.
- Y. Men, Z. Lian, Y. Tang, and J. Xiao, “Dyntypo: Example-based dynamic text effects transfer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 5870–5879.
- C. Barnes, E. Shechtman, A. Finkelstein, and D. B. Goldman, “Patchmatch: A randomized correspondence algorithm for structural image editing,” ACM Trans. Graph., vol. 28, no. 3, p. 24, 2009.
- A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” arXiv preprint arXiv:1706.03762, 2017.
- S. Yang, Z. Wang, Z. Wang, N. Xu, J. Liu, and Z. Guo, “Controllable artistic text style transfer via shape-matching gan,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 4442–4451.
- C. Barnes, E. Shechtman, D. B. Goldman, and A. Finkelstein, “The generalized patchmatch correspondence algorithm,” in European Conference on Computer Vision. Springer, 2010, pp. 29–43.
- S. Yang, J. Liu, W. Wang, and Z. Guo, “Tet-gan: Text effects transfer via stylization and destylization,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, no. 01, 2019, pp. 1238–1245.
- K. S. Bhat, S. M. Seitz, J. K. Hodgins, and P. K. Khosla, “Flow-based video synthesis and editing,” in ACM SIGGRAPH 2004 Papers, 2004, pp. 360–363.
- P. Bénard, F. Cole, M. Kass, I. Mordatch, J. Hegarty, M. S. Senn, K. Fleischer, D. Pesare, and K. Breeden, “Stylizing animation by example,” ACM Transactions on Graphics (TOG), vol. 32, no. 4, pp. 1–12, 2013.
- J. Fišer, O. Jamriška, D. Simons, E. Shechtman, J. Lu, P. Asente, M. Lukáč, and D. Sỳkora, “Example-based synthesis of stylized facial animations,” ACM Transactions on Graphics (TOG), vol. 36, no. 4, pp. 1–11, 2017.
- P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” arXiv preprint arXiv:2012.09841, 2020.
- A. v. d. Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” arXiv preprint arXiv:1711.00937, 2017.
- W. Yan, Y. Zhang, P. Abbeel, and A. Srinivas, “Videogpt: Video generation using vq-vae and transformers,” arXiv preprint arXiv:2104.10157, 2021.
- V. Kwatra, I. Essa, A. Bobick, and N. Kwatra, “Texture optimization for example-based synthesis,” in ACM SIGGRAPH 2005 Papers, 2005, pp. 795–802.
- P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017.
Paper Prompts
Sign up for free to create and run prompts on this paper.