Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

MeetDot: Videoconferencing with Live Translation Captions (2109.09577v1)

Published 20 Sep 2021 in cs.CL and cs.AI

Abstract: We present MeetDot, a videoconferencing system with live translation captions overlaid on screen. The system aims to facilitate conversation between people who speak different languages, thereby reducing communication barriers between multilingual participants. Currently, our system supports speech and captions in 4 languages and combines automatic speech recognition (ASR) and machine translation (MT) in a cascade. We use the re-translation strategy to translate the streamed speech, resulting in caption flicker. Additionally, our system has very strict latency requirements to have acceptable call quality. We implement several features to enhance user experience and reduce their cognitive load, such as smooth scrolling captions and reducing caption flicker. The modular architecture allows us to integrate different ASR and MT services in our backend. Our system provides an integrated evaluation suite to optimize key intrinsic evaluation metrics such as accuracy, latency and erasure. Finally, we present an innovative cross-lingual word-guessing game as an extrinsic evaluation metric to measure end-to-end system performance. We plan to make our system open-source for research purposes.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (8)
  1. Arkady Arkhangorodsky (6 papers)
  2. Christopher Chu (4 papers)
  3. Scot Fang (4 papers)
  4. Yiqi Huang (13 papers)
  5. Denglin Jiang (4 papers)
  6. Ajay Nagesh (7 papers)
  7. Boliang Zhang (9 papers)
  8. Kevin Knight (29 papers)
Citations (3)

Summary

We haven't generated a summary for this paper yet.