Papers
Topics
Authors
Recent
Search
2000 character limit reached

Do VSR Models Generalize Beyond LRS3?

Published 23 Nov 2023 in cs.CV, cs.CL, and cs.LG | (2311.14063v1)

Abstract: The Lip Reading Sentences-3 (LRS3) benchmark has primarily been the focus of intense research in visual speech recognition (VSR) during the last few years. As a result, there is an increased risk of overfitting to its excessively used test set, which is only one hour duration. To alleviate this issue, we build a new VSR test set named WildVSR, by closely following the LRS3 dataset creation processes. We then evaluate and analyse the extent to which the current VSR models generalize to the new test data. We evaluate a broad range of publicly available VSR models and find significant drops in performance on our test set, compared to their corresponding LRS3 results. Our results suggest that the increase in word error rates is caused by the models inability to generalize to slightly harder and in the wild lip sequences than those found in the LRS3 test set. Our new test benchmark is made public in order to enable future research towards more robust VSR models.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (42)
  1. Lrs3-ted: a large-scale dataset for visual speech recognition. arXiv preprint arXiv:1809.00496, 2018.
  2. ASR is all you need: cross-modal distillation for lip reading. ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 2020-May:2143–2147, 11 2019.
  3. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in neural information processing systems, 33:12449–12460, 2020.
  4. Conformers are all you need for visual speech recogntion. arXiv preprint arXiv:2302.10915, 2023.
  5. François Chollet. On the measure of intelligence. arXiv preprint arXiv:1911.01547, 2019.
  6. Voxceleb2: Deep speaker recognition. arXiv preprint arXiv:1806.05622, 2018.
  7. Lip reading in the wild. In Computer Vision–ACCV 2016: 13th Asian Conference on Computer Vision, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part II 13, pages 87–103. Springer, 2017.
  8. Out of time: automated lip sync in the wild. In Computer Vision–ACCV 2016 Workshops: ACCV 2016 International Workshops, Taipei, Taiwan, November 20-24, 2016, Revised Selected Papers, Part II 13, pages 251–263. Springer, 2017.
  9. Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. arXiv preprint arXiv:1804.03619, 2018.
  10. Testing the manifold hypothesis. Journal of the American Mathematical Society, 29(4):983–1049, 2016.
  11. Visual speech-aware perceptual 3d facial expression reconstruction from videos. arXiv preprint arXiv:2207.11094, 2022.
  12. Algorithm of scene change detection in a video sequence based on the threedimensional histogram of color images. In 2014 12th International Conference on Actual Problems of Electronics Instrument Engineering (APEIE), pages 1–1, 2014.
  13. Jointly learning visual and auditory speech representations from raw data. arXiv preprint arXiv:2212.06246, 2022.
  14. A survey on contrastive self-supervised learning. Technologies, 9(1):2, 2020.
  15. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  16. Testing the correlation of word error rate and perplexity. Speech Communication, 38(1-2):19–28, 2002.
  17. Neural substrates participating in acquisition of facial familiarity: an fmri study. Neuroimage, 20(3):1734–1742, 2003.
  18. Learning a model of facial shape and expression from 4D scans. ACM Transactions on Graphics (TOG), 36(6), 11 2017.
  19. Synthvsr: Scaling up visual speech recognition with synthetic supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18806–18815, 2023.
  20. Auto-avsr: Audio-visual speech recognition with automatic labels. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE, 2023.
  21. Lira: Learning visual speech representations from audio through self-supervision. arXiv preprint arXiv:2106.09171, 2021.
  22. End-to-end audio-visual speech recognition with conformers. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7613–7617. IEEE, 2021.
  23. Visual speech recognition for multiple languages in the wild. Nature Machine Intelligence, pages 1–10, 2022.
  24. Training strategies for improved lip-reading. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 8472–8476. IEEE, 2022.
  25. Contrastive Learning of Global-Local Video Representations. Advances in Neural Information Processing Systems, 9:7025–7040, 4 2021.
  26. Recurrent neural network transducer for audio-visual speech recognition. In 2019 IEEE automatic speech recognition and understanding workshop (ASRU), pages 905–912. IEEE, 2019.
  27. Praktische verfahren der gleichungsauflösung. ZAMM-Journal of Applied Mathematics and Mechanics/Zeitschrift für Angewandte Mathematik und Mechanik, 9(1):58–77, 1929.
  28. Sub-word level lip reading with visual attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5162–5172, 2022.
  29. Yolo5face: Why reinventing a face detector. In Computer Vision–ECCV 2022 Workshops: Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part V, pages 228–244. Springer, 2023.
  30. Robust speech recognition via large-scale weak supervision. arXiv preprint arXiv:2212.04356, 2022.
  31. Do imagenet classifiers generalize to imagenet? In International conference on machine learning, pages 5389–5400. PMLR, 2019.
  32. Audio-visual speech recognition is worth 32x32x8 voxels. In 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 796–802. IEEE, 2021.
  33. Lightface: A hybrid deep face recognition framework. In 2020 Innovations in Intelligent Systems and Applications Conference (ASYU), pages 23–27. IEEE, 2020.
  34. Hyperextended lightface: A facial attribute analysis framework. In 2021 International Conference on Engineering and Emerging Technologies (ICEET), pages 1–4. IEEE, 2021.
  35. Cross-modal self-supervised learning for lip reading: When contrastive learning meets adversarial training. In Proceedings of the 29th ACM International Conference on Multimedia, pages 2456–2464, 2021.
  36. Learning audio-visual speech representation by masked multimodal cluster prediction. arXiv preprint arXiv:2201.02184, 2022.
  37. Robust self-supervised audio-visual speech recognition. arXiv preprint arXiv:2201.01763, 2022.
  38. Lip reading sentences in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6447–6456, 2017.
  39. Christopher Summerfield. Natural General Intelligence: How understanding the brain can help us build AI. Oxford University Press, 2022.
  40. Alexander Todorov. The role of the amygdala in face perception and evaluation. Motivation and Emotion, 36:16–26, 2012.
  41. Ledyard R Tucker. Some mathematical notes on three-mode factor analysis. Psychometrika, 31(3):279–311, 1966.
  42. Vatlm: Visual-audio-text pre-training with unified masked prediction for speech representation learning. arXiv preprint arXiv:2211.11275, 2022.
Citations (3)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.