Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

vTTS: visual-text to speech (2203.14725v1)

Published 28 Mar 2022 in cs.SD

Abstract: This paper proposes visual-text to speech (vTTS), a method for synthesizing speech from visual text (i.e., text as an image). Conventional TTS converts phonemes or characters into discrete symbols and synthesizes a speech waveform from them, thus losing the visual features that the characters essentially have. Therefore, our method synthesizes speech not from discrete symbols but from visual text. The proposed vTTS extracts visual features with a convolutional neural network and then generates acoustic features with a non-autoregressive model inspired by FastSpeech2. Experimental results show that 1) vTTS is capable of generating speech with naturalness comparable to or better than a conventional TTS, 2) it can transfer emphasis and emotion attributes in visual text to speech without additional labels and architectures, and 3) it can synthesize more natural and intelligible speech from unseen and rare characters than conventional TTS.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (5)
  1. Yoshifumi Nakano (1 paper)
  2. Takaaki Saeki (22 papers)
  3. Shinnosuke Takamichi (70 papers)
  4. Katsuhito Sudoh (35 papers)
  5. Hiroshi Saruwatari (100 papers)
Citations (3)

Summary

We haven't generated a summary for this paper yet.