Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Deep Hierarchical Feature Sparse Framework for Occluded Person Re-Identification

Published 15 Jan 2024 in cs.CV | (2401.07469v1)

Abstract: Most existing methods tackle the problem of occluded person re-identification (ReID) by utilizing auxiliary models, resulting in a complicated and inefficient ReID framework that is unacceptable for real-time applications. In this work, a speed-up person ReID framework named SUReID is proposed to mitigate occlusion interference while speeding up inference. The SUReID consists of three key components: hierarchical token sparsification (HTS) strategy, non-parametric feature alignment knowledge distillation (NPKD), and noise occlusion data augmentation (NODA). The HTS strategy works by pruning the redundant tokens in the vision transformer to achieve highly effective self-attention computation and eliminate interference from occlusions or background noise. However, the pruned tokens may contain human part features that contaminate the feature representation and degrade the performance. To solve this problem, the NPKD is employed to supervise the HTS strategy, retaining more discriminative tokens and discarding meaningless ones. Furthermore, the NODA is designed to introduce more noisy samples, which further trains the ability of the HTS to disentangle different tokens. Experimental results show that the SUReID achieves superior performance with surprisingly fast inference.

Authors (2)
Definition Search Book Streamline Icon: https://streamlinehq.com
References (41)
  1. Q. Liu, D. Yuan, N. Fan, P. Gao, X. Li and Z. He, “Learning Dual-Level Deep Representation for Thermal Infrared Tracking,” in IEEE Transactions on Multimedia, vol. 25, pp. 1269-1281, 2023.
  2. D. Yuan, X. Shu, Q. Liu, X. Zhang, and Z. He, “Robust thermal infrared tracking via an adaptively multi-feature fusion model,” Neural Comput Appl, vol. 35, pp.1224-1228, 2023.
  3. Y. Sun, L. Zheng, Y. Yang, Q. Tian, and S. Wang, “Beyond part models: Person retrieval with refined part pooling (and a strong convolutional baseline),” in Proc. Eur. Conf. Comput. Vis., 2018, pp. 480–496.
  4. S. He, H. Luo, P. Wang, F. Wang, H. Li, and W. Jiang, “TransReID: Transformer-based object re-identification,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. 2021, pp. 15013–15022.
  5. H. Luo, Y. Gu, X. Liao, S. Lai and W. Jiang, “A strong baseline and batch normalization neck for deep person re-identification,” in IEEE Transactions on Multimedia, vol. 22, no. 10, pp. 2597–2609, 2019.
  6. Y. Sun, C. Cheng, Y. Zhang, C. Zhang, L. Zheng, Z. Wang, and Y. Wei, “Circle loss: A unified perspective of pair similarity optimization,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2020, pp. 6398–6407.
  7. J. Miao, Y. Wu, P. Liu, Y. Ding, and Y. Yang (2019) “Pose-guided feature alignment for occluded person re-identification,” in Proc. IEEE Int. Conf. Comput. Vis., 2019, pp. 542–551.
  8. G. Wang, S. Yang, H. Liu, Z. Wang, Y. Yang, S. Wang, G. Yu, E. Zhou, and J. Sun, (2020) “High-order information matters: Learning relation and topology for occluded person re-identification,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2020, pp. 6449–6458.
  9. S. Gao, J. Wang, H. Lu, and Z. Liu, “Pose-guided visible part matching for occluded person reid,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2020, pp. 11744–11752.
  10. L. He, Y. Wang, W. Liu, H. Zhao, Z. Sun, and J. Feng, “Foreground-aware pyramid reconstruction for alignment-free occluded person re identification,” in Proc. IEEE Int. Conf. Comput. Vis., 2019, pp. 8450– 8459.
  11. K. Zhu, H. Guo, Z. Liu, M. Tang, and J. Wang, “Identity-guided human semantic parsing for person re-identification,” in Proc. Eur. Conf. Comput. Vis., 2020, pp. 346–363.
  12. T. Wang, H. Liu, P. Song, T. Guo, and W. Shi, “Pose-guided feature disentangling for occluded person re-identification based on transformer,” in Proc. AAAI Conf. Artif. Intell., 2022, pp. 2540-2549.
  13. L He and W Liu. “Guided saliency feature learning for person re-identification in crowded scenes”. in Proc. Eur. Conf. Comput. Vis., 2020, pp. 357–373.
  14. Y. Li, J. He, T. Zhang, X. Liu, Y. Zhang, and F. Wu, “Diverse part discovery: Occluded person re-identification with part-aware transformer,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., 2021, pp. 2898–2907.
  15. Z. Wang, F. Zhu, S. Tang, R. Zhao, L. He, and J. Song, “Feature erasing and diffusion network for occluded person re-identification,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2022, pp. 4754 – 4763.
  16. Y. Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C. Hsieh, “Dynamicvit: Efficient vision transformers with dynamic token sparsification,” in Proc. Adv. Neural Inf. Process. Syst., 2021. pp. 13937—13949.
  17. G. Hinton, O. Vinyals and J. Dean, “Distilling the knowledge in a neural network,” arXiv:1503.02531, 2015.
  18. J. H. Cho and B. Hariharan, “On the efficacy of knowledge distillation,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. 2019, pp. 4793-4801.
  19. C. Yang, L. Xie, C. Su, and A. Yuille, “Snapshot distillation: Teacher-student optimization in one generation,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2019, pp. 2854-2863.
  20. D. Chen, J. Mei, Y. Zhang, C. Wang, Z. Wang, Y. Feng, and C. Chen, “Cross-layer distillation with semantic calibration,” in Proc. AAAI Conf. Artif. Intell., 2021, pp. 664-680.
  21. B. Heo, J. Kim, S. Yun, H. Park, N. Kwak, and J. Choi. “A comprehensive overhaul of feature distillation,” in Proc. IEEE/CVF Int. Conf. Comput. Vis., 2019, pp. 1921-1930.
  22. X. Jin, B. Peng, Y. Wu, Yu Liu, J. Liu, D. Liang, J. Yan, and X. Hu. “Knowledge distillation via route constrained optimization,” in Proc. IEEE/CVF Int. Conf. Comput. Vis., 2019, pp. 1345-1354.
  23. A. Romero, N. Ballas, S. Kahou, A. Chassang, C. Gatta, and Y. Bengio, “Fitnets: Hints for thin deep nets,” in Proc. Int. Conf. Learn. Representations, 2015, pp. 1-13.
  24. R. He, S. Sun, J. Yang, S. Bai, and X. Qi. “Knowledge distillation as efficient pre-training: faster convergence, higher data-efficiency, and better transferability,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2022, pp. 9151-9161.
  25. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. Adv. Neural Inf. Process. Syst., 2017, pp. 5998–6008.
  26. Y. Liang, C. Ge, Z. Tong, Y. Song, J. Wang, and P. Xie, “EViT: Expediting vision transformers via token reorganizations,” in Proc. Int. Conf. Learn. Representations, 2022, pp. 1-21.
  27. H. Yin, A. Vahdat, J. M. Alvarez, A. Mallya, J. Kautz and P. Molchanov, “A-ViT: Adaptive tokens for efficient vision transformer,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2022, pp. 10799-10808.
  28. Z. Huang and N. Wang, “Like what you like: Knowledge distill via neuron selectivity transfer,” arXiv:1707.01219, 2017.
  29. X. Jin, B. Peng, Y. Wu, Y. Liu, J. Liu, D. Liang, J. Yan, and X. Hu, “Knowledge distillation via route constrained optimization,” in Proc. IEEE/CVF Int. Conf. Comput. Vis., 2019, pp. 1345-1354.
  30. J. Kim, S. Park, and N. Kwak, “Paraphrasing complex network: Network compression via factor transfer,” arXiv:1802.04977, 2018.
  31. K. Nikos and Z. Sergey, “Paying more attention to attention: improving the performance of convolutional neural networks via attention transfer,” in Proc. Int. Conf. Learn. Representations, 2017, pp. 1-13.
  32. H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jegou, “Training data-efficient image transformers & distillation through attention,” in Proc. Int. Conf. Mach. Learn., 2021, pp. 1-13.
  33. Z. Zhong, L. Zheng, G. Kang, S. Li, and Y. Yang, “Random erasing data augmentation,” in Proc. AAAI Conf. Artif. Intell., 2020, pp. 13001–13008.
  34. M. Jia, X. Cheng, S. Lu and J. Zhang, “Learning disentangled representation implicitly via transformer for occluded person re-identification,” in IEEE Transactions on Multimedia, vol. 25, pp. 1294-1305, 2023.
  35. J. Zhuo, Z. Chen, J. Lai, and G. Wang, “Occluded person re-identification,” in Proc. IEEE Conf. Multimedia Expo, 2018, pp. 1–6.
  36. W. Zheng, X. Li, T. Xiang, S. Liao, J. Lai, and S. Gong, “Partial person re-identification,” in Proc. IEEE Int. Conf. Comput. Vis., 2015, pp. 4678–4686.
  37. L. Zheng, L. Shen, L. Tian, S. Wang, J. Wang, and Q. Tian, “Scalable person re-identification: A benchmark,” in Proc. IEEE Int. Conf. Comput. Vis., 2015, pp. 1116–1124.
  38. Z. Zheng, L. Zheng, and Y. Yang, “Unlabeled samples generated by GAN improve the person re-identification baseline in vitro,” in Proc. IEEE Int. Conf. Comput. Vis., 2017, pp. 3754–3762.
  39. S. Wang, B. Huang, H. Li, G. Qi, D. Tao and Z. Yu, “Key point-aware occlusion suppression and semantic alignment for occluded person re-identification,” Inf. Sci., vol. 606, pp. 669-687, 2022.
  40. K. Sun, B. Xiao, D. Liu and J. Wang, “Deep high-resolution representation learning for human pose estimation,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2019, pp. 5686-5696.
  41. K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., 2016, pp. 770–778.
Citations (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.