Acoustically compatible intermediate speech domains for staged adaptation

Investigate whether spontaneous, prosodically rich speech corpora, such as conversational or podcast-style speech, provide a more acoustically compatible intermediate domain for staged adaptation of Whisper models to Greek singing voice transcription.

Background

The paper evaluates a two-stage adaptation strategy in which Whisper is first fine-tuned on approximately 37.5 hours of read Greek speech from Mozilla Common Voice and then fully fine-tuned on the aligned Greek singing corpus. This strategy substantially benefits Whisper Large-v3 but provides little or no improvement for the smaller models. The authors attribute the limited benefit, particularly for Whisper Small, to the controlled prosody of the read-speech corpus, which may reinforce acoustic patterns that are distant from singing.

The unresolved problem is whether a more expressive intermediate speech domain—specifically spontaneous, prosodically rich speech such as conversational or podcast-style speech—would better bridge the speech-to-singing domain gap and improve staged adaptation for Greek automatic lyric transcription.

References

Future work will investigate whether spontaneous, prosodically rich speech corpora (e.g., conversational or podcast-style speech) provide a more acoustically compatible intermediate domain.

Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation  (2609.11302 - Frangiadaki et al., 10 Sep 2026) in Section 5.3, “Two-Stage Adaptation”