Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
102 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Discriminative Spatial-Semantic VOS Solution: 1st Place Solution for 6th LSVOS (2408.16431v1)

Published 29 Aug 2024 in cs.CV

Abstract: Video object segmentation (VOS) is a crucial task in computer vision, but current VOS methods struggle with complex scenes and prolonged object motions. To address these challenges, the MOSE dataset aims to enhance object recognition and differentiation in complex environments, while the LVOS dataset focuses on segmenting objects exhibiting long-term, intricate movements. This report introduces a discriminative spatial-temporal VOS model that utilizes discriminative object features as query representations. The semantic understanding of spatial-semantic modules enables it to recognize object parts, while salient features highlight more distinctive object characteristics. Our model, trained on extensive VOS datasets, achieved first place (\textbf{80.90\%} $\mathcal{J & F}$) on the test set of the 6th LSVOS challenge in the VOS Track, demonstrating its effectiveness in tackling the aforementioned challenges. The code will be available at \href{https://github.com/yahooo-m/VOS-Solution}{code}.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (6)
  1. Deshui Miao (7 papers)
  2. Yameng Gu (2 papers)
  3. Xin Li (980 papers)
  4. Zhenyu He (57 papers)
  5. Yaowei Wang (149 papers)
  6. Ming-Hsuan Yang (377 papers)
Github Logo Streamline Icon: https://streamlinehq.com