Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Audio-visual Saliency for Omnidirectional Videos (2311.05190v1)

Published 9 Nov 2023 in cs.CV

Abstract: Visual saliency prediction for omnidirectional videos (ODVs) has shown great significance and necessity for omnidirectional videos to help ODV coding, ODV transmission, ODV rendering, etc.. However, most studies only consider visual information for ODV saliency prediction while audio is rarely considered despite its significant influence on the viewing behavior of ODV. This is mainly due to the lack of large-scale audio-visual ODV datasets and corresponding analysis. Thus, in this paper, we first establish the largest audio-visual saliency dataset for omnidirectional videos (AVS-ODV), which comprises the omnidirectional videos, audios, and corresponding captured eye-tracking data for three video sound modalities including mute, mono, and ambisonics. Then we analyze the visual attention behavior of the observers under various omnidirectional audio modalities and visual scenes based on the AVS-ODV dataset. Furthermore, we compare the performance of several state-of-the-art saliency prediction models on the AVS-ODV dataset and construct a new benchmark. Our AVS-ODV datasets and the benchmark will be released to facilitate future research.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (9)
  1. Yuxin Zhu (11 papers)
  2. Xilei Zhu (6 papers)
  3. Huiyu Duan (38 papers)
  4. Jie Li (553 papers)
  5. Kaiwei Zhang (11 papers)
  6. Yucheng Zhu (20 papers)
  7. Li Chen (590 papers)
  8. Xiongkuo Min (139 papers)
  9. Guangtao Zhai (231 papers)
Citations (7)

Summary

We haven't generated a summary for this paper yet.