Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
41 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
41 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding (2401.06397v3)

Published 12 Jan 2024 in cs.CV

Abstract: Vision-language foundation models, represented by Contrastive Language-Image Pre-training (CLIP), have gained increasing attention for jointly understanding both vision and textual tasks. However, existing approaches primarily focus on training models to match global image representations with textual descriptions, thereby overlooking the critical alignment between local regions and corresponding text tokens. This paper extends CLIP with multi-granularity alignment. Notably, we deliberately construct a new dataset comprising pseudo annotations at various levels of granularities, encompassing image-level, region-level as well as pixel-level captions and tags. Accordingly, we develop a Unified Multi-Granularity learning framework, termed UMG-CLIP, which simultaneously empowers the model with versatile perception abilities across different levels of detail. With parameter efficient tuning, UMG-CLIP surpasses current widely used CLIP variants and achieves state-of-the-art performance on diverse image understanding benchmarks, including open-world recognition, retrieval, semantic segmentation, and panoptic segmentation tasks. We believe that UMG-CLIP represents a valuable advancement in vision-language foundation models. The code is available at https://github.com/lygsbw/UMG-CLIP.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (11)
  1. Bowen Shi (82 papers)
  2. Peisen Zhao (11 papers)
  3. Zichen Wang (46 papers)
  4. Yuhang Zhang (64 papers)
  5. Yaoming Wang (6 papers)
  6. Jin Li (365 papers)
  7. Wenrui Dai (35 papers)
  8. Junni Zou (31 papers)
  9. Hongkai Xiong (75 papers)
  10. Qi Tian (314 papers)
  11. Xiaopeng Zhang (100 papers)
Citations (4)
X Twitter Logo Streamline Icon: https://streamlinehq.com