Weak- and zero-shot supervision for egocentric graph quality

Establish whether weakly supervised and zero-shot methods can achieve the quality of densely annotated methods for constructing egocentric interaction and scene graphs.

Background

Expressive graph representations require dense labels for entities, relations, and temporal evolution, but such annotations are scarce in egocentric video. This scarcity has encouraged the use of zero-shot foundation-model-based graph construction and weak supervision.

The survey explicitly leaves unresolved whether these scalable alternatives can match the quality of annotation-based graph construction.

References

The annotations are not there. The expressive graphs the field wants depend on dense supervision---Action Genome's, EASG's---that is expensive in third-person video and barely exists for egocentric. This single fact explains the recent lurch toward zero-shot, foundation-model-driven construction, and whether weak and zero-shot supervision can match annotation-based quality remains unresolved.

Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI  (2608.18671 - Zamani et al., 19 Aug 2026) in Section 9.6, “Open Challenges in Graph-based Reasoning” (Sec. graph-challenges)