Symbolic graph integration with vision-language models
Unify discrete graph-based representations of egocentric interactions with the continuous embeddings and open-ended language interface of vision-language models.
References
First, the symbol--vector gap: a graph is discrete and symbolic, a VLM reasons in continuous embeddings, and uniting them is unsolved.
— Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI
(2608.18671 - Zamani et al., 19 Aug 2026) in Section 10.3, “Graph-enhanced Vision-Language Models” (Sec. future-graph)