Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models (2405.05256v1)

Published 8 May 2024 in cs.CV, cs.AI, and cs.LG

Abstract: Mitigating hallucinations in large vision-LLMs (LVLMs) remains an open problem. Recent benchmarks do not address hallucinations in open-ended free-form responses, which we term "Type I hallucinations". Instead, they focus on hallucinations responding to very specific question formats -- typically a multiple-choice response regarding a particular object or attribute -- which we term "Type II hallucinations". Additionally, such benchmarks often require external API calls to models which are subject to change. In practice, we observe that a reduction in Type II hallucinations does not lead to a reduction in Type I hallucinations but rather that the two forms of hallucinations are often anti-correlated. To address this, we propose THRONE, a novel object-based automatic framework for quantitatively evaluating Type I hallucinations in LVLM free-form outputs. We use public LLMs (LMs) to identify hallucinations in LVLM responses and compute informative metrics. By evaluating a large selection of recent LVLMs using public datasets, we show that an improvement in existing metrics do not lead to a reduction in Type I hallucinations, and that established benchmarks for measuring Type I hallucinations are incomplete. Finally, we provide a simple and effective data augmentation method to reduce Type I and Type II hallucinations as a strong baseline.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (7)
  1. Prannay Kaul (5 papers)
  2. Zhizhong Li (22 papers)
  3. Hao Yang (328 papers)
  4. Yonatan Dukler (10 papers)
  5. Ashwin Swaminathan (18 papers)
  6. C. J. Taylor (4 papers)
  7. Stefano Soatto (179 papers)
Citations (6)

Summary

We haven't generated a summary for this paper yet.