Papers
Topics
Authors
Recent
Search
2000 character limit reached

DARS: Dysarthria-Aware Rhythm-Style Synthesis for ASR Enhancement

Published 2 Mar 2026 in cs.SD and cs.CL | (2603.01369v1)

Abstract: Dysarthric speech exhibits abnormal prosody and significant speaker variability, presenting persistent challenges for automatic speech recognition (ASR). While text-to-speech (TTS)-based data augmentation has shown potential, existing methods often fail to accurately model the pathological rhythm and acoustic style of dysarthric speech. To address this, we propose DARS, a dysarthria-aware rhythm-style synthesis framework based on the Matcha-TTS architecture. DARS incorporates a multi-stage rhythm predictor optimized by contrastive preferences between normal and dysarthric speech, along with a dysarthric-style conditional flow matching mechanism, jointly enhancing temporal rhythm reconstruction and pathological acoustic style simulation. Experiments on the TORGO dataset demonstrate that DARS achieves a Mean Cepstral Distortion (MCD) of 4.29, closely approximating real dysarthric speech. Adapting a Whisper-based ASR system with synthetic dysarthric speech from DARS achieves a 54.22% relative reduction in word error rate (WER) compared to state-of-the-art methods, demonstrating the framework's effectiveness in enhancing recognition performance.

Summary

  • The paper introduces DARS, a Matcha-TTS-based system that combines pause-aware duration prediction, contrastive preference optimization, and global and local style conditioning to synthesize dysarthric speech.
  • DARS reduces mel-cepstral distortion to 4.29 on TORGO, outperforming the Matcha-TTS baseline and showing that all-speaker training can better capture diverse dysarthric patterns than severity-grouped or single-speaker training.
  • Synthetic DARS speech adapts Whisper-Large to real dysarthric speech almost as effectively as real training data, reaching 8.87% overall WER and 7.75% with additional LibriSpeech text, though broader and perceptual evaluations remain necessary.

Motivation and problem statement

Dysarthric speech, produced by speakers with motor speech disorders such as ALS and cerebral palsy, is characterized by slurred articulation, reduced speaking rate, abnormal prosody, and high inter-speaker variability. ASR systems pretrained on typical speech degrade sharply on this population, and the cost of collecting pathological speech data limits the development of robust dysarthric ASR. TTS-based data augmentation is a natural remedy, but existing approaches model the pathological rhythm and acoustic style of dysarthric speech only coarsely: prompt-driven controllable TTS with x-vector speaker adaptation offers limited style control [wagner2025personalized], diffusion-based synthesis lacks precision in rhythm modeling [leung2024training], and severity-conditioned FastSpeech2 with pause insertion struggles with mild dysarthria [soleymanpour2022synthesizing].

Method

DARS builds on Matcha-TTS, a non-autoregressive encoder–decoder TTS system that uses Optimal Transport Conditional Flow Matching (OT-CFM) to learn a vector field from noise to mel-spectrograms, solving the resulting ODE with a first-order Euler step. This backbone was chosen over diffusion probabilistic models to reduce inference cost while retaining automatic text–speech alignment via monotonic alignment search (MAS).

Two mechanisms are added:

Multi-stage rhythm predictor with contrastive preference optimization. A pause predictor first classifies inter-phoneme pauses into KK duration-binned classes (labels derived from Montreal Forced Aligner alignments), trained with cross-entropy; pause embeddings are then inserted into the phoneme sequence, re-encoded by a pause-augmented encoder, and passed to a duration predictor trained with MSE on log-durations against MAS targets. On top of this cascade, a dysarthria-guided CPO loss enforces a margin m=0.75m = 0.75 such that predicted durations lie closer to dysarthric ground-truth durations than to durations predicted by a normal-speech predictor (pretrained on LibriSpeech), with dynamic weights emphasizing pause positions (α≥β\alpha \geq \beta).

Dysarthria-aware acoustic conditional flow matching. The conditional mean μ\bm{\mu} conditioning the flow-matching decoder is augmented with two style representations extracted from reference mel-spectrograms: global style tokens (GST) for global prosody and a frame-level local style encoder with a vector quantization bottleneck, aligned to the pause-augmented hidden states via attention.

Synthesis quality

Experiments use TORGO (8 dysarthric speakers across four severity levels, 7 controls). Three training strategies are compared: All-Speaker (ASp), Dysarthria-Severity-Group (DSpG), and Single-Speaker (SSp). MCD on the validation set improves monotonically with each proposed component:

Model ASp DSpG SSp
Grad-TTS 6.61 6.71 6.81
Matcha-TTS (baseline) 6.25 6.43 6.64
+ rhythm predictor 6.09 6.32 6.57
+ CPO (α=0.7,β=0.3\alpha{=}0.7,\beta{=}0.3) 5.72 5.93 6.08
+ style vectors (full DARS) 4.29 4.46 4.61

The full DARS model reaches an MCD of 4.29 under ASp training, a substantial reduction over both Grad-TTS and the Matcha-TTS baseline, indicating close approximation of real dysarthric acoustic characteristics. ASp consistently outperforms severity-grouped and single-speaker training, which the authors attribute to the larger effective training set — a notable result given the heterogeneity of dysarthric patterns across speakers.

ASR results

Whisper-Large is adapted on synthesized data and evaluated on real TORGO speech, using a 20% evaluation split selected for stability across 10/20/30% samplings (baseline Overall WER ≈ 83.2%). Key findings:

  • Synthetic data nearly matches real data. Full-parameter fine-tuning on ASp-synthesized speech achieves an Overall WER of 8.87%, essentially identical to fine-tuning on real speech (8.85%), an 89.33% relative reduction over the unadapted baseline. LoRA adaptation consistently underperforms full fine-tuning (e.g., 10.12% vs. 8.87% for ASp), suggesting that the distributional gap between pretraining speech and dysarthric speech exceeds what low-rank updates can bridge.
  • Severity-dependent gains. Improvements concentrate in severe categories: relative WER reductions reach 93.54% for moderately-severe speakers, whereas mild speakers start near ceiling.
  • Comparison with prior systems. Against Soleymanpour et al. (FastSpeech2 + DNN-HMM, Overall WER 39.2%) and Leung et al. (Grad-TTS + Whisper, 16.93%), DARS achieves 8.87%. Extending synthesis to LibriSpeech text with TORGO reference audio yields an Overall WER of 7.75%, a 54.22% relative reduction over the strongest prior system, supporting cross-corpus generalization of the augmentation pipeline.

A caveat applies to these comparisons: E18 uses a weaker DNN-HMM recognizer, so part of its gap reflects ASR architecture rather than synthesis quality alone; the more meaningful comparison is E19, which shares the Whisper backbone.

Limitations and open questions

The evaluation is confined to TORGO's eight dysarthric speakers, all with ALS or CP, and the paper does not report subjective listening tests or objective prosodic metrics beyond MCD, leaving the perceptual naturalness of the synthesized rhythm unverified. The claim that synthetic data "closely matches" real speech rests on MCD, whose correlation with subjective quality is asserted rather than demonstrated here. Additionally, the CPO margin and the α/β\alpha/\beta weighting are set empirically without ablation over a broader range, and the mechanism relies on MFA-derived pause labels whose reliability on severely disordered speech is not analyzed. Whether the ASp advantage persists at much larger corpus scales, or whether severity-grouped models become preferable as per-severity data grows, remains open.

Conclusion

DARS couples a pause-then-duration rhythm predictor trained with dysarthria-guided contrastive preferences to a style-conditioned flow matching decoder within Matcha-TTS, reducing MCD to 4.29 on TORGO and enabling Whisper-Large adaptation that attains an Overall WER of 7.75% — a 54.22% relative improvement over the strongest comparable system. The central empirical finding is that jointly modeling pathological rhythm and acoustic style allows purely synthetic data to substitute for real dysarthric speech in ASR adaptation, including out-of-domain text, though validation beyond TORGO and perceptual evaluation remain outstanding questions.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.