A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models (2405.09589v4)

Published 15 May 2024 in cs.LG, cs.AI, cs.CL, cs.CV, cs.SD, and eess.AS

Abstract: The rapid advancement of foundation models (FMs) across language, image, audio, and video domains has shown remarkable capabilities in diverse tasks. However, the proliferation of FMs brings forth a critical challenge: the potential to generate hallucinated outputs, particularly in high-stakes applications. The tendency of foundation models to produce hallucinated content arguably represents the biggest hindrance to their widespread adoption in real-world scenarios, especially in domains where reliability and accuracy are paramount. This survey paper presents a comprehensive overview of recent developments that aim to identify and mitigate the problem of hallucination in FMs, spanning text, image, video, and audio modalities. By synthesizing recent advancements in detecting and mitigating hallucination across various modalities, the paper aims to provide valuable insights for researchers, developers, and practitioners. Essentially, it establishes a clear framework encompassing definition, taxonomy, and detection strategies for addressing hallucination in multimodal foundation models, laying the foundation for future research in this pivotal area.

References (122)

Authors (6)

Pranab Sahoo (5 papers)
Prabhash Meharia (1 paper)
Akash Ghosh (14 papers)
Sriparna Saha (48 papers)
Vinija Jain (43 papers)
Aman Chadha (110 papers)

Citations (4)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Tweets

https://twitter.com/ArxivSound/status/1791318653440311689

https://twitter.com/ArxivSound/status/1792768164561920229

https://twitter.com/AudioAndSpeech/status/1792919214048579830

https://twitter.com/gastronomy/status/1791318675615625502

https://twitter.com/AudioAndSpeech/status/1838743332907806897

https://twitter.com/jkumarsharma998/status/1860937039119835257

A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models (2405.09589v4)

Summary

Related Papers

Tweets