Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
41 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
41 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units (2212.08055v2)

Published 15 Dec 2022 in cs.CL, cs.SD, and eess.AS

Abstract: Direct speech-to-speech translation (S2ST), in which all components can be optimized jointly, is advantageous over cascaded approaches to achieve fast inference with a simplified pipeline. We present a novel two-pass direct S2ST architecture, UnitY, which first generates textual representations and predicts discrete acoustic units subsequently. We enhance the model performance by subword prediction in the first-pass decoder, advanced two-pass decoder architecture design and search strategy, and better training regularization. To leverage large amounts of unlabeled text data, we pre-train the first-pass text decoder based on the self-supervised denoising auto-encoding task. Experimental evaluations on benchmark datasets at various data scales demonstrate that UnitY outperforms a single-pass speech-to-unit translation model by 2.5-4.2 ASR-BLEU with 2.83x decoding speed-up. We show that the proposed methods boost the performance even when predicting spectrogram in the second pass. However, predicting discrete units achieves 2.51x decoding speed-up compared to that case.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (10)
  1. Hirofumi Inaguma (42 papers)
  2. Sravya Popuri (18 papers)
  3. Ilia Kulikov (31 papers)
  4. Peng-Jen Chen (26 papers)
  5. Changhan Wang (46 papers)
  6. Yu-An Chung (33 papers)
  7. Yun Tang (42 papers)
  8. Ann Lee (29 papers)
  9. Shinji Watanabe (416 papers)
  10. Juan Pino (50 papers)
Citations (43)