Papers
Topics
Authors
Recent
Search
2000 character limit reached

An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis

Published 19 Mar 2024 in cs.CL, cs.SD, and eess.AS | (2403.12402v1)

Abstract: Speech LMs are promising for high-quality speech synthesis through in-context learning. A typical speech LM takes discrete semantic units as content and a short utterance as prompt, and synthesizes speech which preserves the content's semantics but mimics the prompt's style. However, there is no systematic understanding on how the synthesized audio is controlled by the prompt and content. In this work, we conduct an empirical study of the widely used autoregressive (AR) and non-autoregressive (NAR) speech LMs and provide insights into the prompt design and content semantic units. Our analysis reveals that heterogeneous and nonstationary prompts hurt the audio quality in contrast to the previous finding that longer prompts always lead to better synthesis. Moreover, we find that the speaker style of the synthesized audio is also affected by the content in addition to the prompt. We further show that semantic units carry rich acoustic information such as pitch, tempo, volume and speech emphasis, which might be leaked from the content to the synthesized audio.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (40)
  1. Audiolm: A language modeling approach to audio generation. IEEE ACM Trans. Audio Speech Lang. Process., 31:2523–2533.
  2. Soundstorm: Efficient parallel audio generation. CoRR, abs/2305.09636.
  3. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  4. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  5. Wavlm: Large-scale self-supervised pre-training for full stack speech processing. IEEE J. Sel. Top. Signal Process., 16(6):1505–1518.
  6. Palm: Scaling language modeling with pathways. CoRR, abs/2204.02311.
  7. w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training. In IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2021, Cartagena, Colombia, December 13-17, 2021, pages 244–250. IEEE.
  8. Seamlessm4t-massively multilingual & multimodal machine translation. CoRR, abs/2308.11596.
  9. Simple and controllable music generation. CoRR, abs/2306.05284.
  10. High fidelity neural audio compression. CoRR, abs/2210.13438.
  11. Polyvoice: Language models for speech to speech translation. CoRR, abs/2306.02982.
  12. Textually pretrained speech language models. CoRR, abs/2305.13009.
  13. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE ACM Trans. Audio Speech Lang. Process., 29:3451–3460.
  14. Make-a-voice: Unified voice synthesis with discrete representation. CoRR, abs/2305.19269.
  15. Text-free prosody-aware generative spoken language modeling. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8666–8681, Dublin, Ireland. Association for Computational Linguistics.
  16. Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  17. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  18. On generative spoken language modeling from raw audio. Transactions of the Association for Computational Linguistics, 9:1336–1354.
  19. Voicebox: Text-guided multilingual universal speech generation at scale. CoRR, abs/2306.15687.
  20. On the utility of self-supervised models for prosody-related tasks. In IEEE Spoken Language Technology Workshop, SLT 2022, Doha, Qatar, January 9-12, 2023, pages 1104–1111. IEEE.
  21. Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks. CoRR, abs/2309.07937.
  22. Speechlmscore: Evaluating speech generation using speech language model. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5.
  23. OpenAI. 2023. GPT-4 technical report. CoRR, abs/2303.08774.
  24. fairseq: A fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), pages 48–53, Minneapolis, Minnesota. Association for Computational Linguistics.
  25. Reproducing whisper-style training using an open-source toolkit and publicly available data. CoRR, abs/2309.13876.
  26. Scaling speech technology to 1, 000+ languages. CoRR, abs/2305.13516.
  27. Robust speech recognition via large-scale weak supervision. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 28492–28518. PMLR.
  28. Audiopalm: A large language model that can speak and listen. CoRR, abs/2306.12925.
  29. Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers. CoRR, abs/2304.09116.
  30. Christian J. Steinmetz and Joshua D. Reiss. 2021. pyloudnorm: A simple yet flexible loudness meter in python. In 150th AES Convention.
  31. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971.
  32. Llama 2: Open foundation and fine-tuned chat models. CoRR, abs/2307.09288.
  33. VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 993–1003, Online. Association for Computational Linguistics.
  34. Neural codec language models are zero-shot text to speech synthesizers. CoRR, abs/2301.02111.
  35. Viola: Unified codec language models for speech recognition, synthesis, and translation. CoRR, abs/2305.16107.
  36. Speechx: Neural codec language model as a versatile speech transformer. CoRR, abs/2308.06873.
  37. Soundstream: An end-to-end neural audio codec. IEEE ACM Trans. Audio Speech Lang. Process., 30:495–507.
  38. OPT: open pre-trained transformer language models. CoRR, abs/2205.01068.
  39. Google USM: scaling automatic speech recognition beyond 100 languages. CoRR, abs/2303.01037.
  40. Speak foreign languages with your own voice: Cross-lingual neural codec language modeling. CoRR, abs/2303.03926.
Citations (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.