Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
41 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
41 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Slow-Fast Auditory Streams For Audio Recognition (2103.03516v1)

Published 5 Mar 2021 in cs.SD, cs.CV, and eess.AS

Abstract: We propose a two-stream convolutional network for audio recognition, that operates on time-frequency spectrogram inputs. Following similar success in visual recognition, we learn Slow-Fast auditory streams with separable convolutions and multi-level lateral connections. The Slow pathway has high channel capacity while the Fast pathway operates at a fine-grained temporal resolution. We showcase the importance of our two-stream proposal on two diverse datasets: VGG-Sound and EPIC-KITCHENS-100, and achieve state-of-the-art results on both.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Evangelos Kazakos (13 papers)
  2. Arsha Nagrani (62 papers)
  3. Andrew Zisserman (248 papers)
  4. Dima Damen (83 papers)
Citations (63)