Persistent object tracking across frames in egocentric videos
Establish methods that enable multimodal foundation models to maintain persistent tracking of objects across frames in egocentric videos, thereby supporting a stable world-state memory rather than purely view-dependent evidence, as required by the Spatial Memory tasks in the SAW (Situated Awareness in the Real World) benchmark.
References
Persistent tracking of objects across frames remains an open challenge across models.
Graphs do not survive ego-motion. The detectors and trackers that populate graph nodes were tuned on framed third-person video, and the head motion, blur, and truncation of first-person capture (Section~\ref{sec:challenges}) corrupt node identity and edge assignment directly. That both SAMJAM and FocusGraph resort to heavy foundation models or optical-flow heuristics just to hold their graphs together across motion is the symptom; robust, lightweight egocentric graph construction is the unmet need.