Scalable General-Purpose Segmentation from Motion Cues

Develop a scalable, general-purpose segmentation model solely from motion cues that learns a universal, transferable object prior rather than remaining restricted to motion-salient regions and instance-level predictions.

Background

Unsupervised video segmentation methods use motion as a supervisory signal for discovering objects, but the paper explains that existing approaches typically focus on moving foregrounds or precise instance masks. As a result, they tend to overfit to motion-salient regions, fail to generalize reliably to static or rarely moving objects, and lack the ability to represent objects at multiple granularities.

The unresolved challenge is to use motion-derived supervision to learn a universal object prior that transfers beyond moving entities and supports general-purpose, multi-granularity segmentation. The proposed MoSA framework is presented as an approach toward this goal, but the paper explicitly characterizes the broader objective as remaining unresolved.

References

Ultimately, building a scalable, general-purpose segmentation model solely from motion cues remains an open challenge.

— Seeing as Humans Do: Learning from Motion to Segment Anything Without Supervision  (2609.39785 - Jian et al., 30 Sep 2026) in Section 2, subsection “Motion-to-Objectness Bootstrapping”