MLLMs-Augmented Visual-Language Representation Learning (2311.18765v3)

Published 30 Nov 2023 in cs.CV, cs.AI, cs.CL, and cs.LG

Abstract: Visual-language pre-training has achieved remarkable success in many multi-modal tasks, largely attributed to the availability of large-scale image-text datasets. In this work, we demonstrate that Multi-modal LLMs (MLLMs) can enhance visual-language representation learning by establishing richer image-text associations for image-text datasets. Our approach is simple, utilizing MLLMs to extend multiple diverse captions for each image. To prevent the bias introduced by MLLMs' hallucinations and monotonous language styles, we propose "text shearing" to maintain the quality and availability of extended captions. In image-text retrieval, without introducing additional training cost, our method consistently obtains 5.6 ~ 35.0 and 16.8 ~ 46.1 improvement on Recall@1 under the fine-tuning and zero-shot settings, respectively. Notably, we obtain zero-shot results that are comparable to fine-tuning on target datasets, which encourages more exploration of the versatile use of MLLMs.

PDF HTML Abstract

Summarize PDF Markdown Bookmark Chat (Pro)

References (70)

Authors (8)

Yanqing Liu (48 papers)
Kai Wang (624 papers)
Wenqi Shao (89 papers)
Ping Luo (340 papers)
Yu Qiao (563 papers)
Mike Zheng Shou (165 papers)
Kaipeng Zhang (73 papers)
Yang You (173 papers)

Citations (10)

View on Semantic Scholar

GitHub

GitHub - lyq312318224/MLLMs-Augmented: The official implementation of 《MLLMs-Augmented Visual-Language Representation Learning》 (31 stars)

Tweets

https://twitter.com/superman_space/status/1759640192225583186

MLLMs-Augmented Visual-Language Representation Learning (2311.18765v3)

Related Papers

GitHub

Tweets