Papers
Topics
Authors
Recent
Search
2000 character limit reached

Cluster-based pruning techniques for audio data

Published 21 Sep 2023 in eess.AS and cs.SD | (2309.11922v1)

Abstract: Deep learning models have become widely adopted in various domains, but their performance heavily relies on a vast amount of data. Datasets often contain a large number of irrelevant or redundant samples, which can lead to computational inefficiencies during the training. In this work, we introduce, for the first time in the context of the audio domain, the k-means clustering as a method for efficient data pruning. K-means clustering provides a way to group similar samples together, allowing the reduction of the size of the dataset while preserving its representative characteristics. As an example, we perform clustering analysis on the keyword spotting (KWS) dataset. We discuss how k-means clustering can significantly reduce the size of audio datasets while maintaining the classification performance across neural networks (NNs) with different architectures. We further comment on the role of scaling analysis in identifying the optimal pruning strategies for a large number of samples. Our studies serve as a proof-of-principle, demonstrating the potential of data selection with distance-based clustering algorithms for the audio domain and highlighting promising research avenues.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (34)
  1. “Dataset pruning: Reducing training data by examining generalization influence,” in The Eleventh International Conference on Learning Representations, 2023.
  2. “Impact of data pruning on machine learning algorithm performance,” arXiv:1901.10539, 2019.
  3. “Beyond neural scaling laws: beating power law scaling via data pruning,” in Advances in Neural Information Processing Systems, 2022.
  4. “Audio classification using braided convolutional neural networks,” IET Signal Processing, vol. 14, no. 7, pp. 448–454, 2020.
  5. “Masked conditional neural networks for sound classification,” Applied Soft Computing, vol. 90, pp. 106073, 2020.
  6. “Keyword spotting system and evaluation of pruning and quantization methods on low-power edge microcontrollers,” arXiv:2208.02765, 2022.
  7. “Spiking neural networks trained with backpropagation for low power neuromorphic implementation of voice activity detection,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 8544–8548.
  8. “Pruning deep neural network models of guitar distortion effects,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 256–264, 2023.
  9. “Sample-size determination methodologies for machine learning in medical imaging research: A systematic review,” Canadian Association of Radiologists Journal, vol. 70, no. 4, pp. 344–353, 2019.
  10. “Small data, big decisions: Model selection in the small-data regime,” in Proceedings of the 37th International Conference on Machine Learning. 2020, ICML’20, JMLR.org.
  11. François Chollet et al., “Keras,” https://keras.io, 2015.
  12. Pete Warden, “Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition,” arXiv:1804.03209, 2018.
  13. “SNIP: SINGLE-SHOT NETWORK PRUNING BASED ON CONNECTION SENSITIVITY,” in International Conference on Learning Representations, 2019.
  14. “Sparse training via boosting pruning plasticity with neuroregeneration,” in Advances in Neural Information Processing Systems, 2021, vol. 34.
  15. “The lottery ticket hypothesis: Finding sparse, trainable neural networks,” in International Conference on Learning Representations, 2019.
  16. “Selection via proxy: Efficient data selection for deep learning,” in International Conference on Learning Representations, 2020.
  17. “Deep learning on a data diet: Finding important examples early in training,” in Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, Eds. 2021, vol. 34, pp. 20596–20607, Curran Associates, Inc.
  18. “Diet selective-backprop: Accelerating training in deep learning by pruning examples,” http://cs231n.stanford.edu/reports/2022/pdfs/93.pdf.
  19. “Prune then distill: Dataset distillation with importance sampling,” in ICASSP 2023, 2023.
  20. “An empirical study of example forgetting during deep neural network learning,” CoRR, vol. abs/1812.05159, 2018.
  21. “Dataset Pruning for Resource-constrained Spoofed Audio Detection,” in Proc. Interspeech 2022, 2022, pp. 416–420.
  22. Vitaly Feldman, “Does Learning Require Memorization? A Short Tale about a Long Tail,” in Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing, New York, NY, USA, 2020, STOC 2020, p. 954–959, Association for Computing Machinery.
  23. “What neural networks memorize and why: Discovering the long tail via influence estimation,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, Eds. 2020, vol. 33, pp. 2881–2891, Curran Associates, Inc.
  24. “Measuring self-supervised representation quality for downstream classification using discriminative features,” 2023.
  25. “wav2vec 2.0: A framework for self-supervised learning of speech representations,” arXiv:2006.11477, 2020.
  26. “From trees to continuous embeddings and back: Hyperbolic hierarchical clustering,” Advances in Neural Information Processing Systems, vol. 33, pp. 15065–15076, 2020.
  27. “Hyperbolic audio source separation,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5.
  28. “Billion-scale similarity search with GPUs,” IEEE Transactions on Big Data, vol. 7, no. 3, pp. 535–547, 2019.
  29. “PEAF: Learnable Power Efficient Analog Acoustic Features for Audio Recognition,” in Proc. Interspeech 2022, 2022, pp. 381–385.
  30. “Tesla: Test-time self-learning with automatic adversarial augmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 20341–20350.
  31. “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998.
  32. “Deep learning scaling is predictable, empirically,” CoRR, vol. abs/1712.00409, 2017.
  33. “Scaling laws for neural language models,” arXiv:2001.08361, 2020.
  34. “SiDi KWS: A Large-Scale Multilingual Dataset for Keyword Spotting,” in Proc. Interspeech 2022, 2022, pp. 4616–4620.
Citations (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.