Weak- and zero-shot supervision for egocentric graph quality
Establish whether weakly supervised and zero-shot methods can achieve the quality of densely annotated methods for constructing egocentric interaction and scene graphs.
References
The annotations are not there. The expressive graphs the field wants depend on dense supervision---Action Genome's, EASG's---that is expensive in third-person video and barely exists for egocentric. This single fact explains the recent lurch toward zero-shot, foundation-model-driven construction, and whether weak and zero-shot supervision can match annotation-based quality remains unresolved.
— Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI
(2608.18671 - Zamani et al., 19 Aug 2026) in Section 9.6, “Open Challenges in Graph-based Reasoning” (Sec. graph-challenges)