Exploring a Unified Attention-Based Pooling Framework for Speaker Verification (1808.07120v1)

Published 21 Aug 2018 in cs.SD and eess.AS

Abstract: The pooling layer is an essential component in the neural network based speaker verification. Most of the current networks in speaker verification use average pooling to derive the utterance-level speaker representations. Average pooling takes every frame as equally important, which is suboptimal since the speaker-discriminant power is different between speech segments. In this paper, we present a unified attention-based pooling framework and combine it with the multi-head attention. Experiments on the Fisher and NIST SRE 2010 dataset show that involving outputs from lower layers to compute the attention weights can outperform average pooling and achieve better results than vanilla attention method. The multi-head attention further improves the performance.

Authors (4)

Yi Liu (543 papers)
Liang He (202 papers)
Weiwei Liu (51 papers)
Jia Liu (369 papers)

Citations (8)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Exploring a Unified Attention-Based Pooling Framework for Speaker Verification (1808.07120v1)

Summary

Related Papers