Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
110 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

What is Point Supervision Worth in Video Instance Segmentation? (2404.01990v1)

Published 1 Apr 2024 in cs.CV

Abstract: Video instance segmentation (VIS) is a challenging vision task that aims to detect, segment, and track objects in videos. Conventional VIS methods rely on densely-annotated object masks which are expensive. We reduce the human annotations to only one point for each object in a video frame during training, and obtain high-quality mask predictions close to fully supervised models. Our proposed training method consists of a class-agnostic proposal generation module to provide rich negative samples and a spatio-temporal point-based matcher to match the object queries with the provided point annotations. Comprehensive experiments on three VIS benchmarks demonstrate competitive performance of the proposed framework, nearly matching fully supervised methods.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (8)
  1. Shuaiyi Huang (12 papers)
  2. De-An Huang (45 papers)
  3. Zhiding Yu (94 papers)
  4. Shiyi Lan (38 papers)
  5. Subhashree Radhakrishnan (7 papers)
  6. Jose M. Alvarez (90 papers)
  7. Abhinav Shrivastava (120 papers)
  8. Anima Anandkumar (236 papers)
Citations (2)

Summary

We haven't generated a summary for this paper yet.

X Twitter Logo Streamline Icon: https://streamlinehq.com