MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response (2309.08730v3)

Published 15 Sep 2023 in eess.AS, cs.AI, cs.CL, cs.MM, and cs.SD

Abstract: LLMs have shown immense potential in multimodal applications, yet the convergence of textual and musical domains remains not well-explored. To address this gap, we present MusiLingo, a novel system for music caption generation and music-related query responses. MusiLingo employs a single projection layer to align music representations from the pre-trained frozen music audio model MERT with a frozen LLM, bridging the gap between music audio and textual contexts. We train it on an extensive music caption dataset and fine-tune it with instructional data. Due to the scarcity of high-quality music Q&A datasets, we created the MusicInstruct (MI) dataset from captions in the MusicCaps datasets, tailored for open-ended music inquiries. Empirical evaluations demonstrate its competitive performance in generating music captions and composing music-related Q&A pairs. Our introduced dataset enables notable advancements beyond previous ones.

References (30)

Authors (8)

Zihao Deng (20 papers)
Yinghao Ma (24 papers)
Yudong Liu (31 papers)
Rongchen Guo (4 papers)
Ge Zhang (170 papers)
Wenhu Chen (134 papers)
Wenhao Huang (98 papers)
Emmanouil Benetos (89 papers)

Citations (13)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Tweets

https://twitter.com/ArxivSound/status/1775373479300202577

https://twitter.com/AudioAndSpeech/status/1775500549900603658

MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response (2309.08730v3)

Summary

Related Papers

Tweets