Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
4 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

ByteSing: A Chinese Singing Voice Synthesis System Using Duration Allocated Encoder-Decoder Acoustic Models and WaveRNN Vocoders (2004.11012v2)

Published 23 Apr 2020 in eess.AS and cs.SD

Abstract: This paper presents ByteSing, a Chinese singing voice synthesis (SVS) system based on duration allocated Tacotron-like acoustic models and WaveRNN neural vocoders. Different from the conventional SVS models, the proposed ByteSing employs Tacotron-like encoder-decoder structures as the acoustic models, in which the CBHG models and recurrent neural networks (RNNs) are explored as encoders and decoders respectively. Meanwhile an auxiliary phoneme duration prediction model is utilized to expand the input sequence, which can enhance the model controllable capacity, model stability and tempo prediction accuracy. WaveRNN neural vocoders are also adopted as neural vocoders to further improve the voice quality of synthesized songs. Both objective and subjective experimental results prove that the SVS method proposed in this paper can produce quite natural, expressive and high-fidelity songs by improving the pitch and spectrogram prediction accuracy and the models using attention mechanism can achieve best performance.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (9)
  1. Yu Gu (220 papers)
  2. Xiang Yin (99 papers)
  3. Yonghui Rao (1 paper)
  4. Yuan Wan (38 papers)
  5. Benlai Tang (10 papers)
  6. Yang Zhang (1132 papers)
  7. Jitong Chen (15 papers)
  8. Yuxuan Wang (239 papers)
  9. Zejun Ma (78 papers)
Citations (69)

Summary

We haven't generated a summary for this paper yet.