Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
5 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Voice Conversion with Conditional SampleRNN (1808.08311v1)

Published 24 Aug 2018 in cs.SD, cs.LG, and eess.AS

Abstract: Here we present a novel approach to conditioning the SampleRNN generative model for voice conversion (VC). Conventional methods for VC modify the perceived speaker identity by converting between source and target acoustic features. Our approach focuses on preserving voice content and depends on the generative network to learn voice style. We first train a multi-speaker SampleRNN model conditioned on linguistic features, pitch contour, and speaker identity using a multi-speaker speech corpus. Voice-converted speech is generated using linguistic features and pitch contour extracted from the source speaker, and the target speaker identity. We demonstrate that our system is capable of many-to-many voice conversion without requiring parallel data, enabling broad applications. Subjective evaluation demonstrates that our approach outperforms conventional VC methods.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (5)
  1. Cong Zhou (39 papers)
  2. Michael Horgan (2 papers)
  3. Vivek Kumar (62 papers)
  4. Cristina Vasco (1 paper)
  5. Dan Darcy (3 papers)
Citations (19)

Summary

We haven't generated a summary for this paper yet.