Additional fine-tuning runs for unstable language adaptation

Determine whether additional fine-tuning runs can improve automatic speech recognition performance for Finnish and Marathi, whose performance substantially worsened under the evaluated fine-tuning procedures.

Background

The paper evaluates simple fine-tuning and full fine-tuning of Whisper-large-v3 for 102 languages. Although fine-tuning improves performance for most languages, the authors report that Finnish and Marathi are notable exceptions for which fine-tuning produces substantially worse results than the zero-shot Whisper baseline.

The authors attribute this outcome tentatively to instability in fine-tuning and explicitly leave unresolved whether running additional fine-tuning configurations could benefit these two languages. This is a concrete unresolved empirical question concerning the robustness of the proposed language-adaptation methods.

References

For a small set of languages, such as Finnish and Marathi, we find that fine-tuning leads to substantially worse performance. We infer that fine-tuning can sometimes be unstable, and leave open the possibility that additional fine-tuning runs could still benefit these languages.

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models  (2609.09554 - Singh et al., 9 Sep 2026) in Section 5, Results, subsection “ASR Results”