2000 character limit reached
Expressive Machine Dubbing Through Phrase-level Cross-lingual Prosody Transfer (2306.11662v2)
Published 20 Jun 2023 in eess.AS
Abstract: Speech generation for machine dubbing adds complexity to conventional Text-To-Speech solutions as the generated output is required to match the expressiveness, emotion and speaking rate of the source content. Capturing and transferring details and variations in prosody is a challenge. We introduce phrase-level cross-lingual prosody transfer for expressive multi-lingual machine dubbing. The proposed phrase-level prosody transfer delivers a significant 6.2% MUSHRA score increase over a baseline with utterance-level global prosody transfer, thereby closing the gap between the baseline and expressive human dubbing by 23.2%, while preserving intelligibility of the synthesised speech.
- Duo Wang (47 papers)
- Mikolaj Babianski (3 papers)
- Giuseppe Coccia (2 papers)
- Patrick Lumban Tobing (20 papers)
- Ravichander Vipperla (6 papers)
- Viacheslav Klimkov (10 papers)
- Vincent Pollet (4 papers)
- Jakub Swiatkowski (4 papers)