Papers
Topics
Authors
Recent
Search
2000 character limit reached

EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos

Published 30 Jul 2024 in cs.CV, cs.MM, cs.SD, and eess.AS | (2407.20592v2)

Abstract: We introduce EgoSonics, a method to generate semantically meaningful and synchronized audio tracks conditioned on silent egocentric videos. Generating audio for silent egocentric videos could open new applications in virtual reality, assistive technologies, or for augmenting existing datasets. Existing work has been limited to domains like speech, music, or impact sounds and cannot capture the broad range of audio frequencies found in egocentric videos. EgoSonics addresses these limitations by building on the strengths of latent diffusion models for conditioned audio synthesis. We first encode and process paired audio-video data to make them suitable for generation. The encoded data is then used to train a model that can generate an audio track that captures the semantics of the input video. Our proposed SyncroNet builds on top of ControlNet to provide control signals that enables generation of temporally synchronized audio. Extensive evaluations and a comprehensive user study show that our model outperforms existing work in audio quality, and in our proposed synchronization evaluation method. Furthermore, we demonstrate downstream applications of our model in improving video summarization.

Authors (2)
Definition Search Book Streamline Icon: https://streamlinehq.com
References (54)
  1. Video summarization using deep neural networks: A survey. Proceedings of the IEEE, 109(11):1838–1863, 2021.
  2. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8, 2023.
  3. The autoencoding variational autoencoder. Advances in Neural Information Processing Systems, 33:15077–15087, 2020.
  4. Novel-view acoustic synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6409–6419, 2023.
  5. Soundingactions: Learning how actions sound from narrated egocentric videos. arXiv preprint arXiv:2404.05206, 2024.
  6. Deep cross-modal audio-visual generation. In Proceedings of the on Thematic Workshops of ACM Multimedia 2017, pages 349–357, 2017.
  7. Diffv2s: Diffusion-based video-to-speech synthesis with vision-guided speaker embedding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7812–7821, 2023.
  8. Spatial audio in virtual reality: A systematic review. In Proceedings of the 25th Symposium on Virtual and Augmented Reality, pages 264–268, 2023.
  9. The epic-kitchens dataset: Collection, challenges and baselines. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 43(11):4125–4141, 2021.
  10. Google DeepMind. Veo. 2024.
  11. Clipsonic: Text-to-audio synthesis with unlabeled videos and pretrained language-vision models. In 2023 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), pages 1–5. IEEE, 2023.
  12. High fidelity neural audio compression. arXiv preprint arXiv:2210.13438, 2022.
  13. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12873–12883, 2021.
  14. Sound-to-imagination: An exploratory study on unsupervised crossmodal translation using diverse audiovisual data. arXiv preprint arXiv:2106.01266, 2021.
  15. Expressive text-to-image generation with rich text. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7545–7556, 2023.
  16. Audio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 776–780. IEEE, 2017.
  17. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18995–19012, 2022.
  18. Photorealistic video generation with diffusion models. arXiv preprint arXiv:2312.06662, 2023.
  19. Cmcgan: A uniform framework for cross-modal visual-audio mutual generation. In Proceedings of the AAAI conference on artificial intelligence, 2018.
  20. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  21. Visagesyntalk: Unseen speaker video-to-speech synthesis via speech-visage feature selection. In European Conference on Computer Vision, pages 452–468. Springer, 2022.
  22. Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models. In International Conference on Machine Learning, pages 13916–13932. PMLR, 2023.
  23. Taming visually guided sound generation. arXiv preprint arXiv:2110.08791, 2021.
  24. Audiogen: Textually guided audio generation. arXiv preprint arXiv:2209.15352, 2022.
  25. Sound-guided semantic image manipulation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3377–3386, 2022.
  26. Learning visual styles from audio-visual associations. In European Conference on Computer Vision, pages 235–252. Springer, 2022.
  27. Gligen: Open-set grounded text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22511–22521, 2023.
  28. Egocentric video-language pretraining. Advances in Neural Information Processing Systems, 35:7575–7586, 2022.
  29. Sora: A review on background, technology, limitations, and opportunities of large vision models. arXiv preprint arXiv:2402.17177, 2024.
  30. Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models. Advances in Neural Information Processing Systems, 36, 2024.
  31. Learning spatial features from audio-visual correspondence in egocentric videos. CVPR, 2024.
  32. Aria digital twin: A new benchmark dataset for egocentric 3d machine perception. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 20133–20143, 2023.
  33. A fast griffin-lim algorithm. In 2013 IEEE workshop on applications of signal processing to audio and acoustics, pages 1–4. IEEE, 2013a.
  34. A fast griffin-lim algorithm. In 2013 IEEE workshop on applications of signal processing to audio and acoustics, pages 1–4. IEEE, 2013b.
  35. Audio-visual floorplan reconstruction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1183–1192, 2021.
  36. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  37. Zero-shot text-to-image generation. In International conference on machine learning, pages 8821–8831. Pmlr, 2021.
  38. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022.
  39. U-net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18, pages 234–241. Springer, 2015.
  40. Moûsai: Efficient text-to-music diffusion models. 2023.
  41. I hear your true colors: Image guided audio generation. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE, 2023.
  42. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792, 2022.
  43. Activerir: Active audio-visual exploration for acoustic environment modeling. arXiv preprint arXiv:2404.16216, 2024.
  44. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020.
  45. Audeo: Audio generation for a silent performance video. Advances in Neural Information Processing Systems, 33:3325–3337, 2020.
  46. Physics-driven diffusion models for impact sound synthesis from videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9749–9759, 2023.
  47. Sound to visual scene generation by audio-to-visual latent alignment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6430–6440, 2023.
  48. Nvae: A deep hierarchical variational autoencoder. Advances in neural information processing systems, 33:19667–19679, 2020.
  49. Music controlnet: Multiple time-varying controls for music generation. arXiv preprint arXiv:2311.07069, 2023.
  50. Multimodal learning with transformers: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023.
  51. Raphael: Text-to-image generation via large mixture of diffusion paths. Advances in Neural Information Processing Systems, 36, 2024.
  52. Reco: Region-controlled text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14246–14255, 2023.
  53. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023.
  54. Self-supervised multimodal learning: A survey. arXiv preprint arXiv:2304.01008, 2023.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 2 tweets with 9 likes about this paper.