If CLIP Could Talk: Understanding Vision-Language Model Representations Through Their Preferred Concept Descriptions (2403.16442v2)

Published 25 Mar 2024 in cs.CL, cs.CV, and cs.LG

Abstract: Recent works often assume that Vision-LLM (VLM) representations are based on visual attributes like shape. However, it is unclear to what extent VLMs prioritize this information to represent concepts. We propose Extract and Explore (EX2), a novel approach to characterize textual features that are important for VLMs. EX2 uses reinforcement learning to align a LLM with VLM preferences and generates descriptions that incorporate features that are important for the VLM. Then, we inspect the descriptions to identify features that contribute to VLM representations. Using EX2, we find that spurious descriptions have a major role in VLM representations despite providing no helpful information, e.g., Click to enlarge photo of CONCEPT. More importantly, among informative descriptions, VLMs rely significantly on non-visual attributes like habitat (e.g., North America) to represent visual concepts. Also, our analysis reveals that different VLMs prioritize different attributes in their representations. Overall, we show that VLMs do not simply match images to scene descriptions and that non-visual or even spurious descriptions significantly influence their representations.

PDF HTML Abstract

Summarize Bookmark Chat (Pro)

References (64)

Authors (3)

Reza Esfandiarpoor (8 papers)
Cristina Menghini (13 papers)
Stephen H. Bach (33 papers)

Citations (3)

View on Semantic Scholar

Tweets

https://twitter.com/stevebach/status/1786048690957726163

If CLIP Could Talk: Understanding Vision-Language Model Representations Through Their Preferred Concept Descriptions (2403.16442v2)

Related Papers

Tweets