Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
51 tokens/sec
GPT-4o
60 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
8 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

What is the Visual Cognition Gap between Humans and Multimodal LLMs? (2406.10424v1)

Published 14 Jun 2024 in cs.CV and cs.AI

Abstract: Recently, Multimodal LLMs (MLLMs) have shown great promise in language-guided perceptual tasks such as recognition, segmentation, and object detection. However, their effectiveness in addressing visual cognition problems that require high-level reasoning is not well-established. One such challenge is abstract visual reasoning (AVR) -- the cognitive ability to discern relationships among patterns in a set of images and extrapolate to predict subsequent patterns. This skill is crucial during the early neurodevelopmental stages of children. Inspired by the AVR tasks in Raven's Progressive Matrices (RPM) and Wechsler Intelligence Scale for Children (WISC), we propose a new dataset MaRs-VQA and a new benchmark VCog-Bench containing three datasets to evaluate the zero-shot AVR capability of MLLMs and compare their performance with existing human intelligent investigation. Our comparative experiments with different open-source and closed-source MLLMs on the VCog-Bench revealed a gap between MLLMs and human intelligence, highlighting the visual cognitive limitations of current MLLMs. We believe that the public release of VCog-Bench, consisting of MaRs-VQA, and the inference pipeline will drive progress toward the next generation of MLLMs with human-like visual cognition abilities.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (8)
  1. Xu Cao (88 papers)
  2. Bolin Lai (23 papers)
  3. Wenqian Ye (24 papers)
  4. Yunsheng Ma (26 papers)
  5. Joerg Heintz (1 paper)
  6. Jintai Chen (57 papers)
  7. Jianguo Cao (10 papers)
  8. James M. Rehg (91 papers)
Citations (5)