ANCHOR: LLM-driven News Subject Conditioning for Text-to-Image Synthesis
Abstract: Text-to-Image (T2I) Synthesis has made tremendous strides in enhancing synthesized image quality, but current datasets evaluate model performance only on descriptive, instruction-based prompts. Real-world news image captions take a more pragmatic approach, providing high-level situational and Named-Entity (NE) information and limited physical object descriptions, making them abstractive. To evaluate the ability of T2I models to capture intended subjects from news captions, we introduce the Abstractive News Captions with High-level cOntext Representation (ANCHOR) dataset, containing 70K+ samples sourced from 5 different news media organizations. With LLMs (LLM) achieving success in language and commonsense reasoning tasks, we explore the ability of different LLMs to identify and understand key subjects from abstractive captions. Our proposed method Subject-Aware Finetuning (SAFE), selects and enhances the representation of key subjects in synthesized images by leveraging LLM-generated subject weights. It also adapts to the domain distribution of news images and captions through custom Domain Fine-tuning, outperforming current T2I baselines on ANCHOR. By launching the ANCHOR dataset, we hope to motivate research in furthering the Natural Language Understanding (NLU) capabilities of T2I models.
- CITE: A corpus of image-text discourse relations. arXiv [cs.CL], April 2019.
- Improving image generation with better captions, 2023.
- Microsoft COCO captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, April 2015.
- VQGAN-CLIP: Open domain image generation and editing with natural language guidance. arXiv:2204.08583 [cs], April 2022.
- ArcFace: Additive angular margin loss for deep face recognition. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, June 2019.
- Retinaface: Single-shot multi-level face localisation in the wild. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 5203–5212. IEEE, 2020.
- CogView: Mastering Text-to-Image generation via transformers. May 2021.
- Stephanie Federico. These are NPR’s photo caption guidelines. NPR, January 2016.
- LayoutGPT: Compositional visual planning and generation with large language models. arXiv [cs.CV], May 2023.
- Generative adversarial nets. In Advances in Neural Information Processing Systems, volume 27. Curran Associates, Inc., 2014.
- Herbert P Grice. Logic and conversation. In Speech acts, pp. 41–58. Brill, 1975.
- CLIPScore: A reference-free evaluation metric for image captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 7514–7528, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics.
- GANs trained by a two Time-Scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017.
- LoRA: Low-rank adaptation of large language models, 2021.
- Multi-Concept customization of Text-to-Image diffusion. December 2022.
- The role of ImageNet classes in fréchet inception distance. March 2022.
- BigDatasetGAN: Synthesizing ImageNet with pixel-wise annotations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 21330–21340, 2022.
- LLM-grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models. ArXiv, abs/2305.13655, May 2023.
- Diffusion model with perceptual loss. arXiv [cs.CV], December 2023.
- Visual news: Benchmark and challenges in news image captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 6761–6771, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics.
- GLIDE: Towards photorealistic image generation and editing with Text-Guided diffusion models. arXiv:2112.10741 [cs], March 2022.
- Pragmatic Issue-Sensitive image captioning. In Findings of the Association for Computational Linguistics: EMNLP 2020, pp. 1924–1938, Online, November 2020. Association for Computational Linguistics.
- Toward verifiable and reproducible human evaluation for text-to-image generation. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 14277–14286. IEEE, June 2023.
- Learning transferable visual models from natural language supervision. arXiv:2103.00020 [cs], February 2021.
- Zero-Shot Text-to-Image generation. February 2021.
- Hierarchical Text-Conditional image generation with CLIP latents. arXiv:2204.06125 [cs], April 2022.
- Nils Reimers. sentence-transformers/all-MiniLM-L6-v2 · hugging face. https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2. Accessed: 2024-4-5.
- High-resolution image synthesis with latent diffusion models. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, June 2022.
- DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, June 2023.
- Photorealistic text-to-image diffusion models with deep language understanding. May 2022.
- FaceNet: A unified embedding for face recognition and clustering. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 815–823. IEEE, June 2015.
- LAION-5B: An open large-scale dataset for training next generation image-text models. October 2022.
- Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2556–2565, Melbourne, Australia, July 2018. Association for Computational Linguistics.
- Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning, pp. 2256–2265. PMLR, June 2015.
- Winoground: Probing vision and language models for visio-linguistic compositionality. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5238–5248, 2022.
- Transform and tell: Entity-aware news image captioning. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 13032–13042, Seattle, WA, USA, 2020. IEEE.
- REL: An entity linker standing on the shoulders of giants. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’20, pp. 2197–2200, New York, NY, USA, July 2020. Association for Computing Machinery.
- Attention is all you need. In I Guyon, U Von Luxburg, S Bengio, H Wallach, R Fergus, S Vishwanathan, and R Garnett (eds.), Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017.
- Context-Aware captions from Context-Agnostic supervision. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1070–1079. IEEE, July 2017.
- Exploring CLIP for assessing the look and feel of images. arXiv [cs.CV], July 2022a.
- Review of large vision models and visual prompt engineering. July 2023.
- DiffusionDB: A large-scale prompt gallery dataset for text-to-image generative models. arXiv:2210. 14896 [cs], 2022b.
- Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis. arXiv [cs.CV], June 2023.
- FastComposer: Tuning-free multi-subject image generation with localized attention. arXiv [cs.CV], May 2023.
- ImageReward: Learning and evaluating human preferences for Text-to-Image generation. April 2023.
- AttnGAN: Fine-Grained text to image generation with attentional generative adversarial networks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1316–1324, Salt Lake City, UT, USA, 2018. IEEE.
- Inserting anybody in diffusion models via celeb basis. June 2023.
- Understanding deep learning requires rethinking generalization. arXiv:1611.03530 [cs], February 2017.
- The unreasonable effectiveness of deep features as a perceptual metric. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 586–595, Salt Lake City, UT, 2018. IEEE.
- Towards Language-Free training for Text-to-Image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 17907–17917. openaccess.thecvf.com, 2022.
- DM-GAN: Dynamic memory generative adversarial networks for Text-To-Image synthesis. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5795–5803, Long Beach, CA, USA, 2019. IEEE.
Paper Prompts
Sign up for free to create and run prompts on this paper.