Papers
Topics
Authors
Recent
Search
2000 character limit reached

Audio-Aware Query-Enhanced Transformers (AuTR)

Updated 27 March 2026
  • The paper introduces AuTR, a multimodal transformer that seeds its decoder with audio-derived queries to focus on sounding objects in video frames.
  • It employs modality-specific encoders and a unified transformer backbone, achieving improved Jaccard and F-scores over previous convolutional fusion methods.
  • The architecture incorporates dynamic convolution in the segmentation head to enhance cross-modal binding and support robust performance in multi-sound and open-set scenarios.

Audio-Aware Query-Enhanced Transformers (AuTR) are a class of multimodal neural architectures specifically designed for the task of audio-visual segmentation (AVS). AVS aims to segment the regions in a video frame that correspond to sounding objects, leveraging both audio and visual cues. AuTR addresses limitations of prior convolutional and early-fusion approaches—namely their restricted receptive fields and inadequate modality fusion—by introducing a unified multimodal transformer framework seeded with audio-derived queries. This architecture enables explicit disambiguation of sounding versus silent (but salient) objects, more robust cross-modal binding, and improved generalization in multi-sound and open-set scenarios (Liu et al., 2023).

1. Architectural Overview

The AuTR framework is organized into three consecutive processing stages: (1) modality-specific encoders for visual and audio streams, (2) a multimodal transformer backbone with both encoder and a query-enhanced decoder, and (3) a lightweight segmentation head based on dynamic convolution.

  • Visual Encoder: A pre-trained ResNet-50 or Pyramid Vision Transformer (PVT-v2) consumes RGB frames v∈R3×H0×W0v \in \mathbb{R}^{3 \times H_0 \times W_0}, yielding three visual feature maps at spatial resolutions fv,â„“t∈RCv×Hℓ×Wâ„“f^t_{v,\ell} \in \mathbb{R}^{C_v \times H_\ell \times W_\ell} for â„“=1,2,3\ell=1,2,3.
  • Audio Encoder: A VGGish backbone (pre-trained on AudioSet) converts T×Ha×WaT \times H_a \times W_a log-mel spectrograms into CaC_a-dimensional embeddings fatf^t_a.
  • Audio–Visual Early Fusion: Both visual and audio embeddings are linearly mapped to a common dimension CavC_{av} for each scale â„“\ell. Multi-head attention fuses these as follows:

Q=WQfv,â„“t,K=WKfat,V=WVfatQ = W_Q f^t_{v,\ell},\quad K = W_K f^t_a,\quad V = W_V f^t_a

Attentionâ„“(fv,â„“t,fat)=softmax(QKTd)V\mathrm{Attention}_\ell(f^t_{v,\ell}, f^t_a) = \mathrm{softmax}\left(\frac{QK^T}{\sqrt{d}}\right)V

This produces three audio–visual fused feature maps fv,ℓt∈RCv×Hℓ×Wℓf^t_{v,\ell} \in \mathbb{R}^{C_v \times H_\ell \times W_\ell}0.

2. Multimodal Transformer Backbone

2.1 Multimodal Encoder

The fused feature maps are spatially flattened, scale-stacked, temporally stacked, and enhanced with fixed 2D positional encodings. They are then processed by a 6-layer deformable transformer encoder, enabling self-attention and cross-scale attention to yield deeply aggregated representations fv,ℓt∈RCv×Hℓ×Wℓf^t_{v,\ell} \in \mathbb{R}^{C_v \times H_\ell \times W_\ell}1.

2.2 Audio-Aware Query-Enhanced Decoder

AuTR displaces traditional DETR-style object queries with fv,ℓt∈RCv×Hℓ×Wℓf^t_{v,\ell} \in \mathbb{R}^{C_v \times H_\ell \times W_\ell}2 learnable vectors fv,ℓt∈RCv×Hℓ×Wℓf^t_{v,\ell} \in \mathbb{R}^{C_v \times H_\ell \times W_\ell}3 (fv,ℓt∈RCv×Hℓ×Wℓf^t_{v,\ell} \in \mathbb{R}^{C_v \times H_\ell \times W_\ell}4), each initialized to fv,ℓt∈RCv×Hℓ×Wℓf^t_{v,\ell} \in \mathbb{R}^{C_v \times H_\ell \times W_\ell}5 (audio embedding plus positional offset). Decoder operation comprises, per layer fv,ℓt∈RCv×Hℓ×Wℓf^t_{v,\ell} \in \mathbb{R}^{C_v \times H_\ell \times W_\ell}6: 1. Self-attention among current queries: fv,ℓt∈RCv×Hℓ×Wℓf^t_{v,\ell} \in \mathbb{R}^{C_v \times H_\ell \times W_\ell}7. 2. Cross-attention using the encoded audio-visual tokens: fv,ℓt∈RCv×Hℓ×Wℓf^t_{v,\ell} \in \mathbb{R}^{C_v \times H_\ell \times W_\ell}8. 3. Feed-forward update with residual: fv,ℓt∈RCv×Hℓ×Wℓf^t_{v,\ell} \in \mathbb{R}^{C_v \times H_\ell \times W_\ell}9.

This explicit seeding of the decoder with audio semantics ensures the attention mechanism focuses on modality-coherent (i.e., sounding) regions.

3. Deep Aggregation and Mask Head

A parallel lightweight decoder, following the Feature Pyramid Network (FPN) paradigm, aggregates the original visual features (â„“=1,2,3\ell=1,2,30), audio tokens (â„“=1,2,3\ell=1,2,31), and transformer encoder output (â„“=1,2,3\ell=1,2,32) via cross-attention and progressive top-down upsampling, generating a multi-modal feature map â„“=1,2,3\ell=1,2,33.

The â„“=1,2,3\ell=1,2,34 decoder output vectors are passed through a small MLP, generating per-query convolution kernels â„“=1,2,3\ell=1,2,35, consisting of â„“=1,2,3\ell=1,2,36 and â„“=1,2,3\ell=1,2,37 kernel weights. This results in:

â„“=1,2,3\ell=1,2,38

where each ℓ=1,2,3\ell=1,2,39 is a low-resolution mask for query T×Ha×WaT \times H_a \times W_a0.

4. Training Objectives and Optimization

For each training frame, only a single ground-truth mask T×Ha×WaT \times H_a \times W_a1 (the sounding object) is provided. A bipartite matching cost is computed:

T×Ha×WaT \times H_a \times W_a2

with T×Ha×WaT \times H_a \times W_a3 (soft Dice loss), T×Ha×WaT \times H_a \times W_a4 (binary focal loss), and T×Ha×WaT \times H_a \times W_a5 (binary cross-entropy for sounding-score head T×Ha×WaT \times H_a \times W_a6).

The total loss is applied only to the matched query:

T×Ha×WaT \times H_a \times W_a7

Typical loss weights: T×Ha×WaT \times H_a \times W_a8, T×Ha×WaT \times H_a \times W_a9, CaC_a0.

Key hyperparameters:

  • CaC_a1 (audio-aware queries)
  • Transformer: 6 encoder and 6 decoder layers, 8 attention heads, CaC_a2 hidden dimension
  • Feature dims: CaC_a3, CaC_a4, CaC_a5
  • AdamW optimizer: learning rate CaC_a6, weight decay CaC_a7; 50 epochs on a single NVIDIA RTX 3090 (batch size 8, modality backbones frozen)

5. Empirical Results

Quantitative performance is assessed on AVSBench S4 (single sound) and MS3 (multi-sound) benchmarks as well as open-set evaluation. On S4 with ResNet-50:

  • TPAVI achieves CaC_a8; AuTR reaches CaC_a9 fatf^t_a0.
  • With PVT-v2: TPAVI at fatf^t_a1, AuTR at fatf^t_a2. For MS3 (multi-sound) without fine-tuning (ResNet-50): TPAVI at fatf^t_a3, AuTR at fatf^t_a4. After S4fatf^t_a5MS3 fine-tuning, TPAVI drops from fatf^t_a6 to fatf^t_a7, while AuTR improves from fatf^t_a8 to fatf^t_a9 CavC_{av}0.

Open-set results (PVT-v2):

  • Seen: AuTR CavC_{av}1, TPAVI CavC_{av}2
  • Unseen: AuTR CavC_{av}3, TPAVI CavC_{av}4

Ablation results:

  • Removing audio-aware queries: Jaccard drops from CavC_{av}5 to CavC_{av}6.
  • Removing dynamic-conv mask head: CavC_{av}7.

Qualitative assessment reveals AuTR's ability to isolate truly sounding objects even in cluttered, multi-object, or unseen-category scenes, whereas TPAVI often selects silent but salient distractors.

6. Distinctive Features and Comparative Analysis

The explicit integration of audio signals as the initial state for transformer queries results in focused and interpretable attention; visual–spatial areas emitting sound are emphasized, while equally salient but soundless objects are suppressed. The combination of dynamic convolution in the segmentation head and deep transformer-based cross-modal fusion distinguishes AuTR from prior convolutional fusion methods.

Earlier fusion-based AVS approaches are limited by small receptive fields and superficial fusion. In contrast, AuTR's transformer stack, multi-scale aggregation, and learned audio query initialization allow for more comprehensive cross-modal reasoning and improved domain adaptation, particularly in multi-sound and open-set regimes.

7. Pseudocode and Workflow Illustration

A schematic summary of AuTR’s per-frame inference loop is as follows: CavC_{av}8

A plausible implication is that further exploration of modality-aware query initialization and adaptive fusion mechanisms could extend AuTR’s robustness to additional domains, such as cross-modal retrieval or active sound-source localization (Liu et al., 2023).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Audio-Aware Query-Enhanced Transformers (AuTR).