Papers
Topics
Authors
Recent
Search
2000 character limit reached

Pac-HuBERT: Self-Supervised Music Separation

Updated 1 July 2026
  • The paper introduces a novel three-stage pipeline that uses primitive auditory feature clustering, masked prediction with a transformer encoder, and a Res-U-Net decoder to address data scarcity in music source separation.
  • The method processes stereo spectrograms through a 2D convolutional encoder and a 12-layer transformer, producing robust embeddings that enhance the separation of vocals, drums, bass, and other sources.
  • Quantitative evaluations on MusDB18 demonstrate that integrating self-supervised pretraining yields significant SDR improvements compared to baseline models, even with limited supervised data.

Pac-HuBERT is a self-supervised learning (SSL) framework for music source separation that leverages large corpora of unlabeled music mixtures to address the long-standing limitation of scarce clean source (stem) data. Drawing inspiration from the HuBERT model for speech representation learning, Pac-HuBERT introduces a three-stage pipeline involving primitive auditory feature clustering, masked prediction pretraining with a transformer-based architecture, and supervised fine-tuning with a Res-U-Net decoder for isolating different musical sources (Chen et al., 2023).

1. Motivation and Architectural Overview

The primary challenge in music source separation is the limited availability of stereo mixes paired with clean source stems, as typified by the MusDB18 dataset (100 training songs, 50 test songs). This data scarcity constrains data-hungry models and motivates leveraging unlabeled music recordings. Pac-HuBERT is constructed to learn representations from vast collections of uncurated mixtures by unsupervised clustering of low-level auditory features and self-supervised prediction of cluster indices, which are termed “primitive auditory clusters.” At separation time, a pretrained encoder (Pac-HuBERT) is connected to a compact Res-U-Net decoder for end-to-end fine-tuning on available labeled data.

The architecture of Pac-HuBERT comprises:

  • A 2D convolutional encoder mapping stereo spectrograms to bottleneck features,
  • A multi-layer transformer encoder (the “bottleneck”) featuring multi-head self-attention and feedforward sublayers,
  • A Res-U-Net decoder with symmetric deconvolution blocks and skip connections, outputting per-source spectrogram masks.

2. Input Representation and Primitive Auditory Feature Extraction

Input audio consists of stereo waveforms sampled at 44.1 kHz and transformed into short-time Fourier transform (STFT) magnitude spectrograms: X[c,t,f]=∑nx[c,n] w[n−tH] e−j2πfn/FX[c,t,f] = \sum_{n} x[c,n]\,w[n - tH]\,e^{-j2\pi fn/F} with window length 2048, hop size 441, yielding T=320T=320 time frames and F=1024F=1024 frequency bins (excluding the Nyquist bin).

Primitive auditory features are computed on each spectrogram using six grouping algorithms that output frame-level foreground/background estimates:

  • Harmonic–Percussive Source Separation (HPSS)
  • REPET and REPET-SIM
  • Two 2D Fourier common-fate techniques (FT2D-M and FT2D-R)
  • Melody-contour extraction (Melodia)

Each of the six algorithms generates two masks (foreground/background), resulting in 12-dimension feature vectors per time-frequency bin.

For subsequent patch-level clustering:

  • Spectrograms are partitioned into non-overlapping time-frequency patches (Pt=32P_t=32, Pf=64P_f=64).
  • Each patch is split into low/high frequency sub-patches.
  • The 12 features are averaged per sub-patch and concatenated, yielding 24 features per patch per channel, or 48 total per stereo patch.

3. Unsupervised Clustering via K-means

The 48-dimensional patch features are extracted from the FMA-Large dataset (106,574 tracks, ∼\sim890 hours) and pooled across the corpus. K-means clustering (Lloyd’s algorithm) is executed with K=960K=960 clusters, and each time-frequency patch in every track is assigned a cluster index c∈{1,…,960}c \in \{1,\ldots,960\}. These indices are used as pseudo-labels for the masked prediction pretraining objective.

4. Self-Supervised Masked-Unit Prediction Pretraining

Pac-HuBERT adapts the HuBERT pretraining objective for time-frequency (TF) music data. The convolutional encoder transforms input spectrograms into bottleneck features: S∈RCb×(T/Pt)×(F/Pf)S \in \mathbb{R}^{C_b \times (T/P_t) \times (F/P_f)} with Cb=384C_b=384. The features are reshaped into a sequential input to a 12-layer transformer encoder (8 heads, hidden size T=320T=3200, feedforward 4T=320T=3201). Each sequence element is mapped by a projection head to an embedding T=320T=3202.

Pretraining employs span masking: 40% of patch positions are randomly chosen as span starts and masked for 5 contiguous patches, with a mask token input to the model at those locations. The objective is to predict the original cluster ID at masked positions using a temperature-scaled cosine-similarity cross-entropy loss: T=320T=3203 where T=320T=3204 is the set of masked positions, T=320T=3205 are learned cluster embeddings, and T=320T=3206.

Pretraining details:

  • Dataset: FMA-Large
  • Optimizer: AdamW, base LR T=320T=3207, 32,000 linear warmup steps, total 250,000 steps
  • Batch size: 96 (8T=320T=3208NVIDIA A40 GPUs)
  • Mask ratio: 40%, span length 5

5. Fine-Tuning with Res-U-Net Decoder and Separation Loss

After pretraining, the Pac-HuBERT encoder-transformer is paired with a Res-U-Net decoder. The decoder consists of six 2D deconvolutional blocks, each upsampling by T=320T=3209, with skip-connections from the corresponding encoder blocks. The network outputs one spectrogram mask per source (vocals, drums, bass, other), which is applied to the input magnitude spectrogram, followed by inverse STFT and overlap-add to reconstruct time-domain waveforms F=1024F=10240.

Fine-tuning is performed on MusDB18 (84 songs train, 16 val), optimizing the L1 loss on reconstructions: F=1024F=10241 Fine-tuning uses:

  • LR F=1024F=10242, similar scheduler as pretraining
  • Batch size: 96
  • 200,000 steps

For a time-domain variant (16 kHz, HuBERT-based), LR F=1024F=10243, batch size 128, and alternate scheduling are applied.

6. Quantitative Evaluation on MusDB18

Median source-to-distortion ratio (SDR, dB) on the MusDB18 test set (50 songs, SI-SNR frame-wise metric) for various models is summarized in the following table.

Model Vocals Drums Bass Other
Demucs V2 (no SSL, 16 kHz) 5.02 6.02 5.40 3.41
Res-U-Net (no SSL) 7.83 5.47 5.21 4.90
Pac-HuBERT (↓ no pretrain) 8.07 5.78 5.21 5.29
Pac-HuBERT (+ SSL pretrain) 8.32 5.86 6.01 5.38
Pac-HuBERT (2× SSL) 8.52 6.20 5.76 5.18

Integrating self-supervised pretraining with Pac-HuBERT yields improvements over baseline models:

  • Compared to Res-U-Net (no SSL pretrain), Pac-HuBERT (+SSL pretrain) achieves +0.49 dB (vocals), +0.39 dB (drums), +0.80 dB (bass), +0.48 dB (other).
  • Pac-HuBERT outperforms Demucs V2 by a wider margin on vocals and "other".
  • With only 25% of the supervised data, SSL pretraining increases vocals SDR by +1.03 dB over the non-pretrained variant.

7. Significance and Broader Implications

Pac-HuBERT demonstrates that self-supervised learning via primitive auditory clustering and masked prediction unlocks significant improvements for music source separation, particularly when annotated data are scarce (Chen et al., 2023). By mining structure from unlabeled music mixtures, Pac-HuBERT delivers transferable representations that amplify the performance of downstream supervised models. This approach suggests broader potential for self-supervised frameworks combining primitive feature clustering with masked modeling in other domains where stem data are limited.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Pac-HuBERT.