Silkie: Preference Distillation for Large Visual Language Models (2312.10665v1)

Published 17 Dec 2023 in cs.CV and cs.CL

Abstract: This paper explores preference distillation for large vision LLMs (LVLMs), improving their ability to generate helpful and faithful responses anchoring the visual context. We first build a vision-language feedback (VLFeedback) dataset utilizing AI annotation. Specifically, responses are generated by models sampled from 12 LVLMs, conditioned on multi-modal instructions sourced from various datasets. We adopt GPT-4V to assess the generated outputs regarding helpfulness, visual faithfulness, and ethical considerations. Furthermore, the preference supervision is distilled into Qwen-VL-Chat through the direct preference optimization (DPO) method. The resulting model Silkie, achieves 6.9% and 9.5% relative improvement on the MME benchmark regarding the perception and cognition capabilities, respectively. Silkie also demonstrates reduced hallucination by setting a new state-of-the-art score of 3.02 on the MMHal-Bench benchmark. Further analysis shows that DPO with our VLFeedback dataset mainly boosts the fine-grained perception and complex cognition abilities of LVLMs, leading to more comprehensive improvements compared to human-annotated preference datasets.

PDF HTML Abstract

Summarize Bookmark Chat (Pro)

Authors (9)

Lei Li (1293 papers)
Zhihui Xie (17 papers)
Mukai Li (17 papers)
Shunian Chen (15 papers)
Peiyi Wang (48 papers)
Liang Chen (360 papers)
Yazheng Yang (16 papers)
Benyou Wang (109 papers)
Lingpeng Kong (134 papers)

Citations (50)

View on Semantic Scholar

Tweets

https://twitter.com/GAIS_jp/status/1743210736422154388

Silkie: Preference Distillation for Large Visual Language Models (2312.10665v1)

Related Papers

Tweets