Papers
Topics
Authors
Recent
Search
2000 character limit reached

ArrayDPS-Refine: Generative Refinement of Discriminative Multi-Channel Speech Enhancement

Published 25 Mar 2026 in eess.AS | (2603.24385v1)

Abstract: Multi-channel speech enhancement aims to recover clean speech from noisy multi-channel recordings. Most deep learning methods employ discriminative training, which can lead to non-linear distortions from regression-based objectives, especially under challenging environmental noise conditions. Inspired by ArrayDPS for unsupervised multi-channel source separation, we introduce ArrayDPS-Refine, a method designed to enhance the outputs of discriminative models using a clean speech diffusion prior. ArrayDPS-Refine is training-free, generative, and array-agnostic. It first estimates the noise spatial covariance matrix (SCM) from the enhanced speech produced by a discriminative model, then uses this estimated noise SCM for diffusion posterior sampling. This approach allows direct refinement of any discriminative model's output without retraining. Our results show that ArrayDPS-Refine consistently improves the performance of various discriminative models, including state-of-the-art waveform and STFT domain models. Audio demos are provided at https://xzwy.github.io/ArrayDPSRefineDemo/.

Summary

  • The paper introduces a training-free, array-agnostic diffusion refinement stage that uses a clean-speech prior and an estimated noise spatial covariance matrix to post-process discriminative multi-channel enhancement outputs.
  • The method improves every tested base model, including TADRN, FaSNet-TAC, and USES2; for example, refined TADRN reaches 0.818 eSTOI, 10.2 dB SI-SDR, and 38.9% WER on the 4-channel test set.
  • The guidance weight provides a practical quality–intelligibility control, but real-world robustness, computational cost, Gaussian-noise assumptions, and automatic parameter selection require further study.

Overview

ArrayDPS-Refine addresses a persistent weakness of discriminative multi-channel speech enhancement: regression-trained models achieve strong objective scores but introduce non-linear distortions that degrade perceptual quality and downstream ASR performance. The paper proposes a training-free, generative, array-agnostic refinement stage that post-processes the output of any discriminative enhancement model using diffusion posterior sampling (DPS) with a pre-trained clean-speech diffusion prior. The method extends ArrayDPS, originally formulated for unsupervised multi-channel source separation under weak white noise, to the enhancement setting by introducing an explicit noise spatial covariance matrix (SCM) estimated from the discriminative model's own output. The authors claim this is the first multi-channel generative method to outperform state-of-the-art discriminative models across perceptual, intelligibility, and WER metrics.

Method

The signal model assumes a single target speaker observed by CC microphones in the STFT domain, Y=H∗ℓX+NY = H *_\ell X + N, with NN modeled as zero-mean complex Gaussian with time-frequency dependent SCM ΦNN(ℓ,k)\Phi_{\text{NN}}(\ell,k). Under this model, the observation likelihood is analytic, enabling DPS-style likelihood guidance.

The pipeline proceeds in two stages. First, given the discriminative model's reference-channel estimate X~=fϕ(Y)\tilde{X} = f_\phi(Y), forward convolutive prediction (FCP) estimates the acoustic transfer function H~\tilde{H}, from which multi-channel reverberant speech and residual noise N~=Y−H~∗ℓX~\tilde{N} = Y - \tilde{H} *_\ell \tilde{X} are reconstructed. The noise SCM is then tracked with a recursive exponential moving average (smoothing factor α=0.95\alpha = 0.95). Second, refinement runs a modified DPS loop initialized at intermediate diffusion step T′=300T' = 300 from the compressed-STFT version of X~\tilde{X} — analogous to StoRM's warm initialization but without requiring joint training. At each reverse step, the one-step Tweedie denoised estimate Y=H∗ℓX+NY = H *_\ell X + N0 is used with FCP (Y=H∗ℓX+NY = H *_\ell X + N1 frames) to re-estimate Y=H∗ℓX+NY = H *_\ell X + N2, form the multi-channel noise residual, and compute a Mahalanobis-type likelihood score weighted by Y=H∗ℓX+NY = H *_\ell X + N3; a guidance scalar Y=H∗ℓX+NY = H *_\ell X + N4 scales this update.

Two design choices are notable. The diffusion prior operates on magnitude-compressed STFT (Y=H∗ℓX+NY = H *_\ell X + N5 phase-preserving compression), following evidence that magnitude compression benefits audio diffusion models. Because the sampled output is not guaranteed to align with the reference channel, a final single-frame FCP filter aligns it with Y=H∗ℓX+NY = H *_\ell X + N6, which is by construction aligned since the discriminative model targets the reference channel.

Experimental setup

Three array-agnostic discriminative baselines are refined: FaSNet-TAC (time-domain, TAC-based), TADRN (triple-path time-domain), and USES2 (STFT-domain SOTA), each also paired with a low-distortion multi-channel Wiener filter (MCWF) baseline. Training data consists of 80,000 simulated 10-second ad-hoc array mixtures (8 microphones within a 0.1 m sphere, 8–16 interfering speakers, 1–50 noise sources, SNR in Y=H∗ℓX+NY = H *_\ell X + N7 dB, SIR in Y=H∗ℓX+NY = H *_\ell X + N8 dB) generated with Pyroomacoustics image-source simulation. The diffusion prior is a 2-D U-Net trained on ~220 hours of DNS-Challenge clean speech for Y=H∗ℓX+NY = H *_\ell X + N9 steps on 8 H100 GPUs. Evaluation uses STOI/eSTOI/WER (Whisper base recognizer), SI-SDR, PESQ, DNSMOS, and UTMOSv2 on 4- and 8-channel test sets.

Results

Refinement improves every metric for every base model, with the largest gains on weaker models:

Model eSTOI (4ch) WB-PESQ (4ch) SI-SDR (4ch) WER % (4ch) UTMOSv2 (4ch)
Noisy 0.361 1.07 −6.3 118 1.82
TADRN 0.783 2.05 8.9 58.4 2.55
Refined TADRN (NN0) 0.818 2.23 10.2 38.9 2.95
FaSNet-TAC 0.677 1.65 5.4 68.9 1.82
Refined FaSNet-TAC (NN1) 0.736 1.88 6.7 56.8 2.59
USES2 0.833 2.42 6.0 36.6 2.90
Refined USES2 (NN2) 0.835 2.41 6.7 32.7 3.07

For TADRN, refinement yields roughly +0.03 eSTOI, +0.2 wideband PESQ, +1 dB SI-SDR, and +0.4 UTMOSv2, while reducing WER by more than 10 points absolute in the 4-channel case and about 15 points in the 8-channel case (e.g., 46.9% → 26.7% at NN3 on 8 channels). FaSNet-TAC gains are larger still (+0.06 eSTOI, +1.3 dB SI-SDR, +0.7 UTMOSv2). Most consequentially, refining USES2 — already superior to TADRN on nearly all metrics — still yields consistent improvement: +0.8 dB SI-SDR, −4 points WER, and +0.17 UTMOSv2 on 4 channels, and +0.5 dB SI-SDR, −6.3 points WER, and +0.1 UTMOSv2 on 8 channels with no metric degradation. This supports the central claim that a generative refiner can improve even SOTA discriminative outputs rather than merely compensating for weak baselines.

The guidance weight NN4 acts as an interpretable trade-off knob: increasing it from 0.4 to 1.2 generally improves intelligibility metrics (STOI, eSTOI, WER) while degrading perceptual metrics (PESQ, DNSMOS, UTMOSv2), reflecting the shift of posterior mass between data prior and mixture consistency. Notably, MCWF baselines built from the same enhanced speech perform substantially worse than both the discriminative models and their refinements, indicating that the diffusion prior contributes information beyond what linear spatial filtering can recover.

Limitations and open questions

Several caveats bear directly on the results. The evaluation relies entirely on simulated ad-hoc array data with image-source reverberation; performance on real recordings with sensor mismatch, non-Gaussian or non-stationary noise, and moving sources is not established. The Gaussian noise assumption underlying the SCM-weighted likelihood is a modeling approximation, and the SCM itself is estimated from the discriminative model's output — so systematic errors in NN5 propagate into the likelihood guidance. The method inherits the computational cost of iterative DPS sampling (300 reverse steps plus per-step FCP estimation), which the paper does not quantify against real-time constraints. Finally, the NN6 trade-off between intelligibility and perceptual quality is set empirically per model; no principled selection rule is offered, and whether a single NN7 generalizes across unseen noise conditions remains open.

Conclusion

ArrayDPS-Refine demonstrates that a training-free diffusion posterior sampler, equipped with an SCM-based likelihood derived from a discriminative model's own output, can consistently refine multi-channel enhancement results — including those of state-of-the-art waveform and STFT-domain models — across intelligibility, quality, and WER metrics. Its array-agnostic formulation makes it applicable to arbitrary microphone configurations without retraining, positioning it as a modular post-processing stage complementary to existing discriminative systems.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.