VL-BEiT: Generative Vision-Language Pretraining (2206.01127v2)

Published 2 Jun 2022 in cs.CV and cs.CL

Abstract: We introduce a vision-language foundation model called VL-BEiT, which is a bidirectional multimodal Transformer learned by generative pretraining. Our minimalist solution conducts masked prediction on both monomodal and multimodal data with a shared Transformer. Specifically, we perform masked vision-LLMing on image-text pairs, masked LLMing on texts, and masked image modeling on images. VL-BEiT is learned from scratch with one unified pretraining task, one shared backbone, and one-stage training. Our method is conceptually simple and empirically effective. Experimental results show that VL-BEiT obtains strong results on various vision-language benchmarks, such as visual question answering, visual reasoning, and image-text retrieval. Moreover, our method learns transferable visual features, achieving competitive performance on image classification, and semantic segmentation.

PDF Abstract

Summarize Bookmark Chat (Pro)

Authors (4)

Hangbo Bao (17 papers)
Wenhui Wang (47 papers)
Li Dong (154 papers)
Furu Wei (291 papers)

Citations (42)

View on Semantic Scholar

YouTube

Show All Videos

VL-BEiT: Generative Vision-Language Pretraining (2206.01127v2)

Related Papers

YouTube