Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
4 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Revitalize Region Feature for Democratizing Video-Language Pre-training of Retrieval (2203.07720v3)

Published 15 Mar 2022 in cs.CV

Abstract: Recent dominant methods for video-language pre-training (VLP) learn transferable representations from the raw pixels in an end-to-end manner to achieve advanced performance on downstream video-language retrieval. Despite the impressive results, VLP research becomes extremely expensive with the need for massive data and a long training time, preventing further explorations. In this work, we revitalize region features of sparsely sampled video clips to significantly reduce both spatial and temporal visual redundancy towards democratizing VLP research at the same time achieving state-of-the-art results. Specifically, to fully explore the potential of region features, we introduce a novel bidirectional region-word alignment regularization that properly optimizes the fine-grained relations between regions and certain words in sentences, eliminating the domain/modality disconnections between pre-extracted region features and text. Extensive results of downstream video-language retrieval tasks on four datasets demonstrate the superiority of our method on both effectiveness and efficiency, \textit{e.g.}, our method achieves competing results with 80\% fewer data and 85\% less pre-training time compared to the most efficient VLP method so far \cite{lei2021less}. The code will be available at \url{https://github.com/showlab/DemoVLP}.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (11)
  1. Guanyu Cai (10 papers)
  2. Yixiao Ge (99 papers)
  3. Binjie Zhang (7 papers)
  4. Alex Jinpeng Wang (20 papers)
  5. Rui Yan (250 papers)
  6. Xudong Lin (37 papers)
  7. Ying Shan (252 papers)
  8. Lianghua He (23 papers)
  9. Xiaohu Qie (22 papers)
  10. Jianping Wu (30 papers)
  11. Mike Zheng Shou (165 papers)
Citations (6)

Summary

We haven't generated a summary for this paper yet.

Github Logo Streamline Icon: https://streamlinehq.com