Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
80 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Towards Multi-Scale Style Control for Expressive Speech Synthesis (2104.03521v1)

Published 8 Apr 2021 in cs.SD, cs.CL, and eess.AS

Abstract: This paper introduces a multi-scale speech style modeling method for end-to-end expressive speech synthesis. The proposed method employs a multi-scale reference encoder to extract both the global-scale utterance-level and the local-scale quasi-phoneme-level style features of the target speech, which are then fed into the speech synthesis model as an extension to the input phoneme sequence. During training time, the multi-scale style model could be jointly trained with the speech synthesis model in an end-to-end fashion. By applying the proposed method to style transfer task, experimental results indicate that the controllability of the multi-scale speech style model and the expressiveness of the synthesized speech are greatly improved. Moreover, by assigning different reference speeches to extraction of style on each scale, the flexibility of the proposed method is further revealed.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (6)
  1. Xiang Li (1002 papers)
  2. Changhe Song (17 papers)
  3. Jingbei Li (12 papers)
  4. Zhiyong Wu (171 papers)
  5. Jia Jia (59 papers)
  6. Helen Meng (204 papers)
Citations (45)