Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
102 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers (2006.10930v2)

Published 19 Jun 2020 in eess.AS, cs.CL, and cs.SD

Abstract: We propose an end-to-end speaker-attributed automatic speech recognition model that unifies speaker counting, speech recognition, and speaker identification on monaural overlapped speech. Our model is built on serialized output training (SOT) with attention-based encoder-decoder, a recently proposed method for recognizing overlapped speech comprising an arbitrary number of speakers. We extend SOT by introducing a speaker inventory as an auxiliary input to produce speaker labels as well as multi-speaker transcriptions. All model parameters are optimized by speaker-attributed maximum mutual information criterion, which represents a joint probability for overlapped speech recognition and speaker identification. Experiments on LibriSpeech corpus show that our proposed method achieves significantly better speaker-attributed word error rate than the baseline that separately performs overlapped speech recognition and speaker identification.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (7)
  1. Naoyuki Kanda (61 papers)
  2. Yashesh Gaur (43 papers)
  3. Xiaofei Wang (138 papers)
  4. Zhong Meng (53 papers)
  5. Zhuo Chen (319 papers)
  6. Tianyan Zhou (11 papers)
  7. Takuya Yoshioka (77 papers)
Citations (73)

Summary

We haven't generated a summary for this paper yet.