Leveraging unlabeled or weakly labeled egocentric video via self-supervised objectives
Investigate whether and how incorporating weaker or unlabeled egocentric human video through self-supervised objectives during pretraining improves the performance and generalization of the EgoScale flow-based Vision–Language–Action policy for dexterous manipulation.
References
Looking forward, several directions remain open. As egocentric human data continues to grow, incorporating weaker or unlabeled video via self-supervised objectives may further amplify these benefits.
The open challenge is scale: these gains depend on detector-driven or LLM-generated annotations whose fidelity is itself limited by the egocentric failure modes of Section~\ref{sec:hoi-limitations}. Closing the verb gap at Ego4D scale, without a dense-annotation crutch, is the work that remains.