Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
5 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation (2104.11710v2)

Published 23 Apr 2021 in cs.SD, cs.CL, and eess.AS

Abstract: The audio segmentation mismatch between training data and those seen at run-time is a major problem in direct speech translation. Indeed, while systems are usually trained on manually segmented corpora, in real use cases they are often presented with continuous audio requiring automatic (and sub-optimal) segmentation. After comparing existing techniques (VAD-based, fixed-length and hybrid segmentation methods), in this paper we propose enhanced hybrid solutions to produce better results without sacrificing latency. Through experiments on different domains and language pairs, we show that our methods outperform all the other techniques, reducing by at least 30% the gap between the traditional VAD-based approach and optimal manual segmentation.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Marco Gaido (47 papers)
  2. Matteo Negri (93 papers)
  3. Mauro Cettolo (20 papers)
  4. Marco Turchi (51 papers)
Citations (23)

Summary

We haven't generated a summary for this paper yet.