Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
4 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer (2407.06524v3)

Published 9 Jul 2024 in cs.SD, cs.MM, and eess.AS

Abstract: Recent speech enhancement methods based on convolutional neural networks (CNNs) and transformer have been demonstrated to efficaciously capture time-frequency (T-F) information on spectrogram. However, the correlation of each channels of speech features is failed to explore. Theoretically, each channel map of speech features obtained by different convolution kernels contains information with different scales demonstrating strong correlations. To fill this gap, we propose a novel dual-branch architecture named channel-aware dual-branch conformer (CADB-Conformer), which effectively explores the long range time and frequency correlations among different channels, respectively, to extract channel relation aware time-frequency information. Ablation studies conducted on DNS-Challenge 2020 dataset demonstrate the importance of channel feature leveraging while showing the significance of channel relation aware T-F information for speech enhancement. Extensive experiments also show that the proposed model achieves superior performance than recent methods with an attractive computational costs.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (5)
  1. Jizhen Li (2 papers)
  2. Xinmeng Xu (17 papers)
  3. Weiping Tu (17 papers)
  4. Yuhong Yang (54 papers)
  5. Rong Zhu (34 papers)

Summary

We haven't generated a summary for this paper yet.