Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Real-Time Audio-Visual End-to-End Speech Enhancement (2303.07005v1)

Published 13 Mar 2023 in eess.AS

Abstract: Audio-visual speech enhancement (AV-SE) methods utilize auxiliary visual cues to enhance speakers' voices. Therefore, technically they should be able to outperform the audio-only speech enhancement (SE) methods. However, there are few works in the literature on an AV-SE system that can work in real time on a CPU. In this paper, we propose a low-latency real-time audio-visual end-to-end enhancement (AV-E3Net) model based on the recently proposed end-to-end enhancement network (E3Net). Our main contribution includes two aspects: 1) We employ a dense connection module to solve the performance degradation caused by the deep model structure. This module significantly improves the model's performance on the AV-SE task. 2) We propose a multi-stage gating-and-summation (GS) fusion module to merge audio and visual cues. Our results show that the proposed model provides better perceptual quality and intelligibility than the baseline E3net model with a negligible computational cost increase.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (6)
  1. Zirun Zhu (8 papers)
  2. Hemin Yang (7 papers)
  3. Min Tang (80 papers)
  4. Ziyi Yang (77 papers)
  5. Sefik Emre Eskimez (28 papers)
  6. Huaming Wang (23 papers)
Citations (4)

Summary

We haven't generated a summary for this paper yet.