Papers
Topics
Authors
Recent
Search
2000 character limit reached

Distilled Non-Semantic Speech Embeddings with Binary Neural Networks for Low-Resource Devices

Published 12 Jul 2022 in cs.SD, cs.LG, and eess.AS | (2207.05784v4)

Abstract: This work introduces BRILLsson, a novel binary neural network-based representation learning model for a broad range of non-semantic speech tasks. We train the model with knowledge distillation from a large and real-valued TRILLsson model with only a fraction of the dataset used to train TRILLsson. The resulting BRILLsson models are only 2MB in size with a latency less than 8ms, making them suitable for deployment in low-resource devices such as wearables. We evaluate BRILLsson on eight benchmark tasks (including but not limited to spoken language identification, emotion recognition, health condition diagnosis, and keyword spotting), and demonstrate that our proposed ultra-light and low-latency models perform as well as large-scale models.

Authors (2)
Definition Search Book Streamline Icon: https://streamlinehq.com
References (26)
  1. Meliusnet: An improved network architecture for binary neural networks. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1439–1448, 2021.
  2. Back to simplicity: How to train accurate bnns from scratch? arXiv preprint arXiv:1906.08637, 2019.
  3. Crema-d: Crowd-sourced emotional multimodal actors dataset. IEEE transactions on affective computing, 5(4):377–390, 2014.
  4. Distilled binary neural network for monaural speech separation. In 2018 International Joint Conference on Neural Networks (IJCNN), pages 1–8. IEEE, 2018.
  5. Binaryconnect: Training deep neural networks with binary weights during propagations. Advances in neural information processing systems, 28, 2015.
  6. Larq: An open-source library for training binarized neural networks. Journal of Open Source Software, 5(45):1746, January 2020.
  7. Vocalsound: A dataset for improving human vocal sounds recognition. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 151–155, 2022.
  8. Distilling the knowledge in a neural network. In NIPS Deep Learning and Representation Learning Workshop, 2015.
  9. Libri-light: A benchmark for asr with limited or no supervision. In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020.
  10. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015.
  11. Training binary neural networks with knowledge transfer. Neurocomputing, 396:534–541, 2020.
  12. Ken MacLean. Voxforge. Ken MacLean.[Online]. Available: http://www. voxforge. org/home.[Acedido em 2012], 2018.
  13. Multilingual spoken words corpus. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021.
  14. FRILL: A Non-Semantic Speech Embedding for Mobile Devices. In Proc. Interspeech 2021, pages 1204–1208, 2021.
  15. Karol J. Piczak. ESC: Dataset for Environmental Sound Classification. In Proceedings of the 23rd Annual ACM Conference on Multimedia, pages 1015–1018. ACM Press, 2015.
  16. Fitnets: Hints for thin deep nets. In In Proceedings of ICLR, 2015.
  17. Contrastive learning of general-purpose audio representations. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 3875–3879. IEEE, 2021.
  18. Leveraging recent advances in deep learning for audio-visual emotion recognition. Pattern Recognition Letters, 146:1–7, 2021.
  19. Universal paralinguistic speech representations using self-supervised conformers. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 3169–3173. IEEE, 2022.
  20. Towards Learning a Universal Non-Semantic Representation of Speech. In Proc. Interspeech 2020, pages 140–144, 2020.
  21. TRILLsson: Distilled Universal Paralinguistic Speech Representations. In Proc. Interspeech 2022, pages 356–360, 2022.
  22. Musan: A music, speech, and noise corpus. arXiv preprint arXiv:1510.08484, 2015.
  23. Efficientnetv2: Smaller models and faster training. In International Conference on Machine Learning, pages 10096–10106. PMLR, 2021.
  24. Pete Warden. Speech commands: A dataset for limited-vocabulary speech recognition. arXiv preprint arXiv:1804.03209, 2018.
  25. Improving knowledge distillation using unified ensembles of specialized teachers. Pattern Recognition Letters, 146:215–221, 2021.
  26. Bigssl: Exploring the frontier of large-scale semi-supervised learning for automatic speech recognition. IEEE Journal of Selected Topics in Signal Processing, 16(6):1519–1532, 2022.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.