dMel: Speech Tokenization made Simple (2407.15835v2)

Published 22 Jul 2024 in cs.CL, cs.AI, cs.SD, and eess.AS

Abstract: LLMs have revolutionized natural language processing by leveraging self-supervised pretraining on vast textual data. Inspired by this success, researchers have investigated complicated speech tokenization methods to discretize continuous speech signals so that LLMing techniques can be applied to speech data. However, existing approaches either model semantic (content) tokens, potentially losing acoustic information, or model acoustic tokens, risking the loss of semantic (content) information. Having multiple token types also complicates the architecture and requires additional pretraining. Here we show that discretizing mel-filterbank channels into discrete intensity bins produces a simple representation (dMel), that performs better than other existing speech tokenization methods. Using an LM-style transformer architecture for speech-text modeling, we comprehensively evaluate different speech tokenization methods on speech recognition (ASR) and speech synthesis (TTS). Our results demonstrate the effectiveness of dMel in achieving high performance on both tasks within a unified framework, paving the way for efficient and effective joint modeling of speech and text.

PDF Abstract

Summarize Bookmark Chat (Pro)

References (40)

Authors (6)

He Bai (50 papers)
Tatiana Likhomanenko (41 papers)
Ruixiang Zhang (69 papers)
Zijin Gu (9 papers)
Zakaria Aldeneh (20 papers)
Navdeep Jaitly (67 papers)

Citations (3)

View on Semantic Scholar

Tweets

https://twitter.com/gm8xx8/status/1815578997360926829

https://twitter.com/knishimae0531/status/1815912236315009114

https://twitter.com/mctalentowen/status/1817817032077062224

dMel: Speech Tokenization made Simple (2407.15835v2)

Related Papers

Tweets