S.T.A.R.-Track: Latent Motion Models for End-to-End 3D Object Tracking with Adaptive Spatio-Temporal Appearance Representations
Abstract: Following the tracking-by-attention paradigm, this paper introduces an object-centric, transformer-based framework for tracking in 3D. Traditional model-based tracking approaches incorporate the geometric effect of object- and ego motion between frames with a geometric motion model. Inspired by this, we propose S.T.A.R.-Track, which uses a novel latent motion model (LMM) to additionally adjust object queries to account for changes in viewing direction and lighting conditions directly in the latent space, while still modeling the geometric motion explicitly. Combined with a novel learnable track embedding that aids in modeling the existence probability of tracks, this results in a generic tracking framework that can be integrated with any query-based detector. Extensive experiments on the nuScenes benchmark demonstrate the benefits of our approach, showing state-of-the-art performance for DETR3D-based trackers while drastically reducing the number of identity switches of tracks at the same time.
- Y. Li, Y. Chen, X. Qi, Z. Li, J. Sun, and J. Jia, “Unifying Voxel-based Representation with Transformer for 3D Object Detection,” in NIPS, 2022.
- C. Badue, R. Guidolini, R. V. Carneiro, P. Azevedo, V. B. Cardoso, A. Forechi, L. Jesus, R. Berriel, T. M. Paixao, F. Mutz, et al., “Self-Driving Cars: A Survey,” Expert Systems with Applications, 2021.
- A. Ess, K. Schindler, B. Leibe, and L. Van Gool, “Object Detection and Tracking for Autonomous Navigation in Dynamic Environments,” IJRR, 2010.
- Y. Wang, V. C. Guizilini, T. Zhang, Y. Wang, H. Zhao, and J. Solomon, “DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries,” in CoRL, 2022.
- S. Doll, R. Schulz, L. Schneider, V. Benzin, M. Enzweiler, and H. P. Lensch, “SpatialDETR: Robust Scalable Transformer-Based 3D Object Detection from Multi-View Camera Images with Global Cross-Sensor Attention,” in ECCV, 2022.
- S. Wang, X. Jiang, and Y. Li, “Focal-PETR: Embracing Foreground for Efficient Multi-Camera 3D Object Detection,” arXiv.org, vol. arXiv:2212.05505, 2022.
- “nuScenes Tracking Task,” https://nuscenes.org/tracking, accessed: 22.02.23.
- P. Karkus, B. Ivanovic, S. Mannor, and M. Pavone, “DiffStack: A Differentiable and Modular Control Stack for Autonomous Vehicles,” in CoRL, 2022.
- J. Gu, C. Hu, T. Zhang, X. Chen, Y. Wang, Y. Wang, and H. Zhao, “ViP3D: End-to-end Visual Trajectory Prediction via 3D Agent Queries,” CVPR, 2023.
- T. Meinhardt, A. Kirillov, L. Leal-Taixe, and C. Feichtenhofer, “TrackFormer: Multi-Object Tracking with Transformers,” in CVPR, 2022.
- F. Zeng, B. Dong, Y. Zhang, T. Wang, X. Zhang, and Y. Wei, “MOTR: End-to-End Multiple-Object Tracking with Transformer,” in ECCV, 2022.
- T. Zhang, X. Chen, Y. Wang, Y. Wang, and H. Zhao, “MUTR3D: A Multi-camera Tracking Framework via 3D-to-2D Queries,” in CVPR Workshops, 2022.
- A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in NIPS, 2017.
- Z. Pang, Z. Li, and N. Wang, “SimpleTrack: Understanding and Rethinking 3D Multi-object Tracking,” in ECCV, 2023.
- F. Ruppel, F. Faion, C. Gläser, and K. Dietmayer, “Transformers for Multi-Object Tracking on Point Clouds,” in IV, 2022.
- H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuScenes: A multimodal dataset for autonomous driving,” in CVPR, 2020.
- N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-End Object Detection with Transformers,” in ECCV, 2020.
- X. Bai, Z. Hu, X. Zhu, Q. Huang, Y. Chen, H. Fu, and C.-L. Tai, “TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers,” in CVPR, 2022.
- G. K. Erabati and H. Araujo, “Li3DeTr: A LiDAR based 3D Detection Transformer,” in WACV, 2023.
- T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common Objects in Context,” in ECCV, 2014.
- T. Yin, X. Zhou, and P. Krahenbuhl, “Center-based 3D Object Detection and Tracking,” in CVPR, 2021.
- J. Pang, L. Qiu, X. Li, H. Chen, Q. Li, T. Darrell, and F. Yu, “Quasi-Dense Similarity Learning for Multiple Object Tracking,” in CVPR, 2021.
- E. Ristani and C. Tomasi, “Features for Multi-Target Multi-Camera Tracking and Re-Identification,” in CVPR, 2018.
- R. Schubert, E. Richter, and G. Wanielik, “Comparison and evaluation of advanced motion models for vehicle tracking,” in IEEE International Conference on Information Fusion, 2008.
- P. Sun, J. Cao, Y. Jiang, R. Zhang, E. Xie, Z. Yuan, C. Wang, and P. Luo, “TransTrack: Multiple Object Tracking with Transformer,” arXiv.org, vol. arXiv:2012.15460, 2020.
- D. Ha, A. Dai, and Q. V. Le, “HyperNetworks,” ICLR, 2017.
- R. E. Kalman, “A New Approach to Linear Filtering and Prediction Problems,” Transactions of the ASME–Journal of Basic Engineering, 1960.
- W. Luo, J. Xing, A. Milan, X. Zhang, W. Liu, and T.-K. Kim, “Multiple Object Tracking: A Literature Review,” AI, 2021.
- G. Ciaparrone, F. L. Sánchez, S. Tabik, L. Troiano, R. Tagliaferri, and F. Herrera, “Deep learning in video multi-object tracking: A survey,” Neurocomputing, 2020.
- Y. Zhou, C. Barnes, J. Lu, J. Yang, and H. Li, “On the Continuity of Rotation Representations in Neural Networks,” in CVPR, 2019.
- T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal Loss for Dense Object Detection,” in CVPR, 2017.
- K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in CVPR, 2016.
- I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” ICLR, 2019.
- Z. Pang, Z. Li, and N. Wang, “Simpletrack: Understanding and rethinking 3d multi-object tracking,” in European Conference on Computer Vision. Springer, 2022, pp. 680–696.
- “MUTR3D Evaluation Github issue #15,” https://github.com/a1600012888/MUTR3D/issues/15.
- T. Fischer, Y.-H. Yang, S. Kumar, M. Sun, and F. Yu, “Cc-3dt: Panoramic 3d object tracking via cross-camera fusion,” CoRL, vol. arXiv:2212.01247, 2022.
- Z. Pang, J. Li, P. Tokmakov, D. Chen, S. Zagoruyko, and Y.-X. Wang, “Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object Tracking,” arXiv.org, vol. arXiv:2302.03802, 2023.
- Y. Liu, T. Wang, X. Zhang, and J. Sun, “PETR: Position Embedding Transformation for Multi-View 3D Object Detection,” in ECCV, 2022.
- Y. Lee, J.-w. Hwang, S. Lee, Y. Bae, and J. Park, “An Energy and GPU-Computation Efficient Backbone Network for Real-Time Object Detection,” in CVPR Workshops, 2019.
- S. Wang, Y. Liu, T. Wang, Y. Li, and X. Zhang, “Exploring object-centric temporal modeling for efficient multi-view 3d object detection,” arXiv.org, vol. arXiv:2303.11926, 2023.
- C. Kim, F. Li, A. Ciptadi, and J. M. Rehg, “Multiple Hypothesis Tracking Revisited,” in ICCV, 2015.
- D. Jia, Y. Yuan, H. He, X. Wu, H. Yu, W. Lin, L. Sun, C. Zhang, and H. Hu, “DETRs with Hybrid Matching,” CVPR, 2023.
Paper Prompts
Sign up for free to create and run prompts on this paper.