Acoustically compatible intermediate speech domains for staged adaptation
Investigate whether spontaneous, prosodically rich speech corpora, such as conversational or podcast-style speech, provide a more acoustically compatible intermediate domain for staged adaptation of Whisper models to Greek singing voice transcription.
References
Future work will investigate whether spontaneous, prosodically rich speech corpora (e.g., conversational or podcast-style speech) provide a more acoustically compatible intermediate domain.
— Automatic Lyric Transcription for Greek Songs: Scaling and Task Composition Effects in Whisper Adaptation
(2609.11302 - Frangiadaki et al., 10 Sep 2026) in Section 5.3, “Two-Stage Adaptation”