Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
119 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching (2406.00320v3)

Published 1 Jun 2024 in cs.SD, cs.CV, cs.MM, and eess.AS

Abstract: Video-to-audio (V2A) generation aims to synthesize content-matching audio from silent video, and it remains challenging to build V2A models with high generation quality, efficiency, and visual-audio temporal synchrony. We propose Frieren, a V2A model based on rectified flow matching. Frieren regresses the conditional transport vector field from noise to spectrogram latent with straight paths and conducts sampling by solving ODE, outperforming autoregressive and score-based models in terms of audio quality. By employing a non-autoregressive vector field estimator based on a feed-forward transformer and channel-level cross-modal feature fusion with strong temporal alignment, our model generates audio that is highly synchronized with the input video. Furthermore, through reflow and one-step distillation with guided vector field, our model can generate decent audio in a few, or even only one sampling step. Experiments indicate that Frieren achieves state-of-the-art performance in both generation quality and temporal alignment on VGGSound, with alignment accuracy reaching 97.22%, and 6.2% improvement in inception score over the strong diffusion-based baseline. Audio samples are available at http://frieren-v2a.github.io.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (8)
  1. Yongqi Wang (24 papers)
  2. Wenxiang Guo (6 papers)
  3. Rongjie Huang (62 papers)
  4. Jiawei Huang (60 papers)
  5. Zehan Wang (38 papers)
  6. Fuming You (6 papers)
  7. Ruiqi Li (44 papers)
  8. Zhou Zhao (219 papers)
Citations (5)

Summary

We haven't generated a summary for this paper yet.