Papers
Topics
Authors
Recent
Search
2000 character limit reached

GCC-PHAT Data Augmentation for SSL

Updated 2 February 2026
  • GDA is a feature-level augmentation strategy that generates synthetic SSL samples by shifting and scaling the dominant GCC-PHAT peaks to rebalance underrepresented classes.
  • The method extracts key peak statistics—positions and amplitudes—from cross-correlation signals to maintain realistic DoA characteristics in augmented data.
  • Empirical results show that GDA, especially when combined with ADIR, enhances SSL accuracy (up to 89%) and reduces mean absolute error, validating its effectiveness in incremental learning.

GCC-PHAT-based Data Augmentation (GDA) is a technique designed to address intra-task class imbalance in sound source localization (SSL) by generating synthetic training features derived from the peak statistics of the Generalized Cross-Correlation with Phase Transform (GCC-PHAT) input domain. GDA operates at the feature level, utilizing the statistical properties of the dominant GCC-PHAT peaks to synthesize new examples for underrepresented (tail) classes, thereby ameliorating the long-tailed direction-of-arrival (DoA) distribution that commonly impairs SSL performance in real-world, incremental learning scenarios (Fan et al., 26 Jan 2026).

1. Mathematical Formulation of GCC-PHAT in SSL

GCC-PHAT transforms the time-domain microphone signals x(t),y(t)x(t), y(t) into frequency domain representations X(f),Y(f)X(f), Y(f). The core feature, the GCC-PHAT function,

Rxy(τ)=∫−∞∞X(f)Y∗(f)∣X(f)Y∗(f)∣ej2πfτdf,R_{xy}(\tau) = \int_{-\infty}^{\infty} \frac{X(f)Y^*(f)}{|X(f)Y^*(f)|} e^{j2\pi f \tau} df,

computes a phase-weighted cross-correlation that accentuates the time-lag corresponding to the DoA between paired microphones. For PP microphone pairs and DaD_a delay bins per pair, feature extraction over a frame yields the vector

x=[Rxy(1)(τ1),…,Rxy(P)(τDa)]⊤∈RP⋅Da.\mathbf{x} = [R^{(1)}_{xy}(\tau_1),\ldots, R^{(P)}_{xy}(\tau_{D_a})]^\top \in \mathbb{R}^{P\cdot D_a}.

These features are the canonical input for downstream incremental SSL models in the described learning pipeline.

2. Peak Statistic Extraction and Characterization

Empirical analysis demonstrates that for each microphone pair, a single dominant peak in the GCC-PHAT domain conveys the DoA information. For each input sample ii and microphone-pair kk (k=1,…,Pk = 1,\ldots,P), GDA extracts:

  • Peak position: pi,k=arg⁡max⁡τRxy(k)(τ)p_{i,k} = \arg\max_{\tau} R^{(k)}_{xy}(\tau),
  • Peak amplitude: X(f),Y(f)X(f), Y(f)0.

Aggregate statistics over all samples of class X(f),Y(f)X(f), Y(f)1 are computed: X(f),Y(f)X(f), Y(f)2 together with empirical variances. No clustering beyond this dominant peak is necessary, as only the highest peak per segment is relocated during augmentation.

3. GDA Algorithmic Procedure

The augmentation process targets all “tail” classes per task X(f),Y(f)X(f), Y(f)3—those with sample count X(f),Y(f)X(f), Y(f)4, where X(f),Y(f)X(f), Y(f)5 and X(f),Y(f)X(f), Y(f)6. For each such class X(f),Y(f)X(f), Y(f)7, X(f),Y(f)X(f), Y(f)8 synthetic examples are generated according to Algorithm 1:

  1. Initialization: Start with X(f),Y(f)X(f), Y(f)9.
  2. For each microphone pair Rxy(τ)=∫−∞∞X(f)Y∗(f)∣X(f)Y∗(f)∣ej2πfτdf,R_{xy}(\tau) = \int_{-\infty}^{\infty} \frac{X(f)Y^*(f)}{|X(f)Y^*(f)|} e^{j2\pi f \tau} df,0:
    • Extract the Rxy(τ)=∫−∞∞X(f)Y∗(f)∣X(f)Y∗(f)∣ej2πfτdf,R_{xy}(\tau) = \int_{-\infty}^{\infty} \frac{X(f)Y^*(f)}{|X(f)Y^*(f)|} e^{j2\pi f \tau} df,1th slice Rxy(τ)=∫−∞∞X(f)Y∗(f)∣X(f)Y∗(f)∣ej2πfτdf,R_{xy}(\tau) = \int_{-\infty}^{\infty} \frac{X(f)Y^*(f)}{|X(f)Y^*(f)|} e^{j2\pi f \tau} df,2 from a base abundant-class sample.
    • Compute shift Rxy(τ)=∫−∞∞X(f)Y∗(f)∣X(f)Y∗(f)∣ej2πfτdf,R_{xy}(\tau) = \int_{-\infty}^{\infty} \frac{X(f)Y^*(f)}{|X(f)Y^*(f)|} e^{j2\pi f \tau} df,3.
    • Cyclically shift Rxy(τ)=∫−∞∞X(f)Y∗(f)∣X(f)Y∗(f)∣ej2πfτdf,R_{xy}(\tau) = \int_{-\infty}^{\infty} \frac{X(f)Y^*(f)}{|X(f)Y^*(f)|} e^{j2\pi f \tau} df,4 by Rxy(τ)=∫−∞∞X(f)Y∗(f)∣X(f)Y∗(f)∣ej2πfτdf,R_{xy}(\tau) = \int_{-\infty}^{\infty} \frac{X(f)Y^*(f)}{|X(f)Y^*(f)|} e^{j2\pi f \tau} df,5 samples along delay.
    • Rescale such that Rxy(τ)=∫−∞∞X(f)Y∗(f)∣X(f)Y∗(f)∣ej2πfτdf,R_{xy}(\tau) = \int_{-\infty}^{\infty} \frac{X(f)Y^*(f)}{|X(f)Y^*(f)|} e^{j2\pi f \tau} df,6.
    • Inject i.i.d. Gaussian noise Rxy(τ)=∫−∞∞X(f)Y∗(f)∣X(f)Y∗(f)∣ej2πfτdf,R_{xy}(\tau) = \int_{-\infty}^{\infty} \frac{X(f)Y^*(f)}{|X(f)Y^*(f)|} e^{j2\pi f \tau} df,7, Rxy(τ)=∫−∞∞X(f)Y∗(f)∣X(f)Y∗(f)∣ej2πfτdf,R_{xy}(\tau) = \int_{-\infty}^{\infty} \frac{X(f)Y^*(f)}{|X(f)Y^*(f)|} e^{j2\pi f \tau} df,8.
    • Place result in Rxy(τ)=∫−∞∞X(f)Y∗(f)∣X(f)Y∗(f)∣ej2πfτdf,R_{xy}(\tau) = \int_{-\infty}^{\infty} \frac{X(f)Y^*(f)}{|X(f)Y^*(f)|} e^{j2\pi f \tau} df,9.
  3. Repeat for PP0 synthetic variants using independent bases.

This strategy ensures synthetic features for tail classes inherit the dominant temporal and amplitude signature of abundant classes, adjusted to the mean DoA characteristics of the target class.

4. Hyperparameterization and Statistical Rationale

Key hyperparameters are:

  • PP1 microphone pairs,
  • PP2 delay bins per pair,
  • PP3: target tail class cardinality reaches 50% of the largest class post-augmentation,
  • PP4: low-variance noise introduces moderate feature diversity.

No secondary amplitude thresholding or multi-peak selection is required. GDA targets exactly the classes beneath the thresholded occupancy, achieving a more uniform class histogram and ensuring rare DoA labels are sufficiently represented.

5. Integration into Incremental SSL Training

In each incremental learning task, prior to classifier update:

  • GCC-PHAT features are extracted from task samples.
  • Tail classes are detected by occupancy.
  • For each, PP5 synthetic features are generated following Algorithm 1 and labeled identically to real data.
  • The data for task PP6 becomes PP7.

The model architecture consists of a frozen 3-layer MLP feature extractor (fixed after first task) and a task-adaptive linear classifier, updated with the Analytic Dynamic Imbalance Rectifier (ADIR). The cross-entropy objective

PP8

is evaluated on the GDA-augmented dataset. No additional regularization is applied. GDA's effect is to increase rare class counts, directly influencing the class-weighting scheme in ADIR (PP9) and reducing global Gini imbalance.

6. Empirical Impact and Ablation

Ablation studies isolate the contribution of GDA in combination and separately from ADIR. Reported results on the SSLR benchmark are:

  • Baseline (no GDA/ADIR): MAE = 7.5°, ACC = 72.0%, BWT = –17.7
  • GDA only: MAE = 7.4°, ACC = 75.0%, BWT = –15.8
  • ADIR only: MAE = 6.1°, ACC = 82.4%, BWT = +1.4
  • GDA + ADIR: MAE = 5.3°, ACC = 89.0%, BWT = +1.6

This demonstrates that GDA independently yields a 3-point increase in accuracy by mitigating intra-task skew. In synergy with ADIR, the framework delivers substantial gains—accuracy increases to 89% and mean absolute error drops to 5.3°, alongside positive backward transfer (Fan et al., 26 Jan 2026).

7. Context and Significance in Robust SSL

GDA addresses structural challenges in SSL arising from unbalanced DoA class histograms, prevalent in naturalistic, evolving sound field deployments. By operating on analytically defined feature statistics and integrating seamlessly with analytic incremental imbalance rectification (ADIR), GDA enables on-the-fly, low-complexity rebalancing of each task dataset. The method avoids the need for storage of past task exemplars and avoids introducing hand-tuned regularizers or high-variance heuristic augmentations. A plausible implication is streamlined adoption for real-world incrementally learned acoustic localization systems, particularly under non-stationary directional statistics.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to GCC-PHAT-based Data Augmentation (GDA).