CAT: Contrastive Adapter Training for Personalized Image Generation
Abstract: The emergence of various adapters, including Low-Rank Adaptation (LoRA) applied from the field of natural language processing, has allowed diffusion models to personalize image generation at a low cost. However, due to the various challenges including limited datasets and shortage of regularization and computation resources, adapter training often results in unsatisfactory outcomes, leading to the corruption of the backbone model's prior knowledge. One of the well known phenomena is the loss of diversity in object generation, especially within the same class which leads to generating almost identical objects with minor variations. This poses challenges in generation capabilities. To solve this issue, we present Contrastive Adapter Training (CAT), a simple yet effective strategy to enhance adapter training through the application of CAT loss. Our approach facilitates the preservation of the base model's original knowledge when the model initiates adapters. Furthermore, we introduce the Knowledge Preservation Score (KPS) to evaluate CAT's ability to keep the former information. We qualitatively and quantitatively compare CAT's improvement. Finally, we mention the possibility of CAT in the aspects of multi-concept adapter and optimization.
- Haokun Liu et al. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. 2022.
- Neil Houlsby et al. Parameter-efficient transfer learning for nlp. 2019.
- Mou et al. T2i- adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. 2023.
- Zhang et al. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. 2023.
- Mahabadi et al. Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks. 2021.
- Robin Rombach et al. High-resolution image synthesis with latent diffusion models. 2022.
- J Sohl-Dickstein. Deep unsupervised learning using nonequilibrium thermodynamics. ICML 2015, 18(1):234–778, 2015.
- Denoising diffusion probabilistic models. arxiv:2006.11239, 2020.
- Alex Nichol Prafulla Dhariwal. Diffusion models beat gans on image synthesis. arXiv:2105.05233, 2021.
- CSaloni Dash et al. Medical time-series data generation using generative adversarial networks. 2020.
- M. AbdulRazek and M. Belal G. Khoriba. Gan-ga: A generative model based on genetic algorithm for medical image generation. 2023.
- Yichun Shi et al. Mvdream: Multi-view diffusion for 3d generation. 2024.
- Hansheng Chen et al. Single-stage diffusion nerf: A unified approach to 3d generation and reconstruction. 2023.
- Jiatao Gu et al. Nerfdiff: Single-image view synthesis with nerf-guided distillation from 3d-aware diffusion. 2023.
- Hu Ye et al. Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. CVPR, 2023.
- Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 22560–22570, October 2023.
- Yoad Tewel et al. Training-free consistent text-to-image generation. https://arxiv.org/abs/2402.03286, 2023.
- Edward J Hu et al. Lora: Low-rank adaptation of large language models. 2022.
- Nataniel Ruiz Hu et al. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. 2023.
- Avrahami et al. The chosen one: Consistent characters in text-to-image diffusion models. 2023.
- Rinon Gal et al. An image is worth one word: Personalizing text-to-image generation using textual inversion. 2022.
- Olaf Ronneberger et al. U-net: Convolutional networks for biomedical image segmentation. 2015.
- Ashish Vaswani et al. Attention is all you need. 2017.
- Alec Radford et al. Learning transferable visual models from natural language supervision. 2021.
- Z Zhao et al. Modified generative adversarial networks for image classification. 2021.
- Hugo Touvron et al. Llama: Open and efficient foundation language models. 2022.
- Tom B. Brown et al. Language models are few-shot learners. 2020.
- Alec Radford et al. Language models are unsupervised multitask learners. 2019.
- Jacob Devlin et al. Bert: Pre-training of deep bidirectional transformers for language understanding. 2019.
- k-means++: the advantages of careful seeding. 2019.
- Peft: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft, 2022.
- Omri Avrahami et al. Break-a-scene: Extracting multiple concepts from a single image. SIGGRAPH Asia 2023, 2023.
- Shin-Ying Yeh et al. Navigating text-to-image customization:from lycoris fine-tuning to model evaluation. 2024.
- Winoground: Probing vision and language models for visio-linguistic compositionality. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
- When and why vision-language models behave like bags-of-words, and what to do about it? In International Conference on Learning Representations, 2023.
- Kihyuk Sohn et al. Styledrop: Text-to-image generation in any style. https://arxiv.org/abs/2306.00983, 2023.
- Zhe Kong et al. Omg: Occlusion-friendly personalized multi-concept generation in diffusion models. https://arxiv.org/abs/2403.10983, 2024.
- Multi-concept customization of text-to-image diffusion. In CVPR, 2023.
- Analyzing appearance and contour based methods for object categorization. In 2003 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2003. Proceedings., volume 2, pages II–409. IEEE, 2003.
- Ilya Loshchilov et al. Decoupled weight decay regularization. ICLR, 2019.
- Dustin Podell et al. Sdxl: Improving latent diffusion models for high-resolution image synthesis. https://arxiv.org/abs/2307.01952, 2023.
- Controlling text-to-image diffusion by orthogonal finetuning. In NeurIPS, 2023.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.