Speech Translation with Large Language Models: An Industrial Practice (2312.13585v1)

Published 21 Dec 2023 in cs.CL, cs.SD, and eess.AS

Abstract: Given the great success of LLMs across various tasks, in this paper, we introduce LLM-ST, a novel and effective speech translation model constructed upon a pre-trained LLM. By integrating the LLM with a speech encoder and employing multi-task instruction tuning, LLM-ST can produce accurate timestamped transcriptions and translations, even from long audio inputs. Furthermore, our findings indicate that the implementation of Chain-of-Thought (CoT) prompting can yield advantages in the context of LLM-ST. Through rigorous experimentation on English and Chinese datasets, we showcase the exceptional performance of LLM-ST, establishing a new benchmark in the field of speech translation. Demo: https://speechtranslation.github.io/LLM-st/.

PDF HTML Abstract

Summarize PDF Markdown Bookmark Chat (Pro)

References (34)

Authors (7)

Zhichao Huang (17 papers)
Rong Ye (20 papers)
Tom Ko (31 papers)
Qianqian Dong (19 papers)
Shanbo Cheng (23 papers)
Mingxuan Wang (83 papers)
Hang Li (277 papers)

Citations (11)

View on Semantic Scholar

GitHub

LLM-ST

Speech Translation with Large Language Models: An Industrial Practice (2312.13585v1)

Related Papers

GitHub