Papers
Topics
Authors
Recent
Search
2000 character limit reached

Denoising-Contrastive Alignment for Continuous Sign Language Recognition

Published 5 May 2023 in cs.CV | (2305.03614v5)

Abstract: Continuous sign language recognition (CSLR) aims to recognize signs in untrimmed sign language videos to textual glosses. A key challenge of CSLR is achieving effective cross-modality alignment between video and gloss sequences to enhance video representation. However, current cross-modality alignment paradigms often neglect the role of textual grammar to guide the video representation in learning global temporal context, which adversely affects recognition performance. To tackle this limitation, we propose a Denoising-Contrastive Alignment (DCA) paradigm. DCA creatively leverages textual grammar to enhance video representations through two complementary approaches: modeling the instance correspondence between signs and glosses from a discrimination perspective and aligning their global context from a generative perspective. Specifically, DCA accomplishes flexible instance-level correspondence between signs and glosses using a contrastive loss. Building on this, DCA models global context alignment between the video and gloss sequences by denoising the gloss representation from noise, guided by video representation. Additionally, DCA introduces gradient modulation to optimize the alignment and recognition gradients, ensuring a more effective learning process. By integrating gloss-wise and global context knowledge, DCA significantly enhances video representations for CSLR tasks. Experimental results across public benchmarks validate the effectiveness of DCA and confirm its video representation enhancement feasibility.

Authors (3)
Definition Search Book Streamline Icon: https://streamlinehq.com
References (47)
  1. Multimodal Machine Learning: A Survey and Taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence 41, 2 (2019), 423–443.
  2. One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale. In ICML, Vol. 202. 1692–1717.
  3. Label-Efficient Semantic Segmentation with Diffusion Models. In ICLR.
  4. Neural sign language translation. In CVPR.
  5. Sign language transformers: Joint end-to-end sign language recognition and translation. In CVPR.
  6. MSDN: Mutually Semantic Distillation Network for Zero-Shot Learning. CVPR, 7602–7611.
  7. Deconstructing Denoising Diffusion Models for Self-Supervised Learning. ArXiv.
  8. A Simple Multi-Modality Transfer Learning Baseline for Sign Language Translation. In CVPR. https://doi.org/10.1109/CVPR52688.2022.00506
  9. Two-Stream Network for Sign Language Recognition and Translation. In NIPS.
  10. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In ICML.
  11. A Kernel Two-Sample Test. J. Mach. Learn. Res. 13 (2012), 723–773.
  12. Distilling Cross-Temporal Contexts for Continuous Sign Language Recognition. In CVPR.
  13. Self-Mutual Distillation Learning for Continuous Sign Language Recognition. In ICCV.
  14. Masked Autoencoders Are Scalable Vision Learners. In CVPR. 15979–15988. https://doi.org/10.1109/CVPR52688.2022.01553
  15. Deep residual learning for image recognition. In CVPR.
  16. Denoising Diffusion Probabilistic Models. In NIPS.
  17. Continuous Sign Language Recognition With Correlation Network. In CVPR.
  18. Self-Emphasizing Network for Continuous Sign Language Recognition. In AAAI.
  19. DiffDis: Empowering Generative Diffusion Model with Cross-Modal Discrimination Capability. ICCV, 15667–15677.
  20. A Survey on Contrastive Self-Supervised Learning. Technologies 9 (2021).
  21. CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language Recognition. In ICCV.
  22. Diederik P Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In ICLR.
  23. Continuous sign language recognition: Towards large vocabulary statistical recognition systems handling multiple signers. Computer Vision and Image Understanding (2015).
  24. UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning. In ACL.
  25. Zekang Liu Lianyu Hu, Liqing Gao and Wei Feng. 2022. Temporal Lift Pooling for Continuous Sign Language Recognition. In ECCV.
  26. Deep Transfer Learning with Joint Adaptation Networks. In ICML, Vol. 70. 2208–2217.
  27. Thop: Pytorch-opcounter. https://github.com/Lyken17/pytorch-OpCounter,2019.8
  28. Visual Alignment Constraint for Continuous Sign Language Recognition. In ICCV.
  29. Deep Radial Embedding for Visual Sequence Learning. In ECCV.
  30. Boosting continuous sign language recognition via cross modality augmentation. In ACMMM.
  31. Iterative alignment network for continuous sign language recognition. In CVPR.
  32. Contrast with reconstruct: Contrastive 3d representation learning guided by generative pretraining. In ICML. 28223–28243.
  33. Learning Transferable Visual Models From Natural Language Supervision. In ICML.
  34. High-Resolution Image Synthesis with Latent Diffusion Models. CVPR (2022), 10674–10685.
  35. MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generation. ArXiv abs/2212.09478 (2022).
  36. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In ICCV. https://doi.org/10.1109/ICCV.2017.74
  37. Denoising Diffusion Implicit Models. In ICLR.
  38. Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation. In ICML.
  39. PAC-Bayes Information Bottleneck. In ICLR.
  40. SimMIM: A Simple Framework for Masked Image Modeling. In CVPR. 9653–9663.
  41. Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models. In CVPR.
  42. Continuous Sign Language Recognition for Hearing-Impaired Consumer Communication via Self-Guidance Network. TCE (2023). https://doi.org/10.1109/TCE.2023.3342163
  43. Alleviating data insufficiency for Chinese sign language recognition. Visual Intelligence 1 (2023), 1–9.
  44. C2ST: Cross-Modal Contextualized Sequence Transduction for Continuous Sign Language Recognition. In ICCV.
  45. CVT-SLR: Contrastive Visual-Textual Transformation for Sign Language Recognition With Variational Alignment. In CVPR.
  46. Improving Sign Language Translation with Monolingual Data by Sign Back-Translation. In CVPR.
  47. Ronglai Zuo and Brian Mak. 2022. C2SLR: Consistency-Enhanced Continuous Sign Language Recognition. In CVPR.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 2 tweets with 0 likes about this paper.