Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model (2406.19905v2)

Published 28 Jun 2024 in cs.CV

Abstract: The Mixture-of-Experts (MoE) has gained increasing attention in studying Large Vision-LLMs (LVLMs). It uses a sparse model to replace the dense model, achieving comparable performance while activating fewer parameters during inference, thus significantly reducing the inference cost. Existing MoE methods in LVLMs encourage different experts to handle different tokens, and they usually employ a router to predict the routing of each token. However, the predictions are based solely on sample features and do not truly reveal the optimization directions of tokens. This may lead to severe optimization interference between different tokens assigned to an expert. To address this problem, this paper proposes a novel method based on token-level gradient analysis, i.e., Solving Token Gradient Conflict (STGC). Specifically, we first use token-level gradients to identify conflicting tokens in experts. After that, we add a specialized loss tailored to eliminate conflicts among tokens within each expert. Our method can serve as a plug-in for diverse Large Vision-LLMs, and extensive experimental results demonstrate its effectiveness. The code will be publicly available at https://github.com/longrongyang/STGC.

PDF HTML Abstract

Summarize PDF Markdown Bookmark Chat (Pro)

Authors (7)

Longrong Yang (3 papers)
Chaoxiang Cai (1 paper)
Fan Yang (877 papers)
Size Li (8 papers)
Di Zhang (230 papers)
Xi Li (197 papers)
Dong Shen (14 papers)

Citations (1)

View on Semantic Scholar

GitHub

GitHub - longrongyang/STGC: Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model (7 stars)

Tweets

https://twitter.com/CSVisionPapers/status/1807851443279249842

Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model (2406.19905v2)

Related Papers

GitHub

Tweets