Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
5 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

What You Say Is What You Show: Visual Narration Detection in Instructional Videos (2301.02307v2)

Published 5 Jan 2023 in cs.CV

Abstract: Narrated ''how-to'' videos have emerged as a promising data source for a wide range of learning problems, from learning visual representations to training robot policies. However, this data is extremely noisy, as the narrations do not always describe the actions demonstrated in the video. To address this problem we introduce the novel task of visual narration detection, which entails determining whether a narration is visually depicted by the actions in the video. We propose What You Say is What You Show (WYS2), a method that leverages multi-modal cues and pseudo-labeling to learn to detect visual narrations with only weakly labeled data. Our model successfully detects visual narrations in in-the-wild videos, outperforming strong baselines, and we demonstrate its impact for state-of-the-art summarization and temporal alignment of instructional videos.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Kumar Ashutosh (17 papers)
  2. Rohit Girdhar (43 papers)
  3. Lorenzo Torresani (73 papers)
  4. Kristen Grauman (136 papers)
Citations (4)

Summary

We haven't generated a summary for this paper yet.