Higher ASR proportions or annealing curricula in multitask fine-tuning

Investigate whether increasing the proportion of aligned speech-text ASR examples beyond 50% or using an annealing curriculum that increases the ASR proportion toward the end of training yields stronger performance in the BuzzASR multitask fine-tuning procedure.

Background

The BuzzASR full fine-tuning pipeline combines aligned speech-text ASR examples with text-only examples. In the reported experiments, a 50% ASR proportion is the best-performing setting among the tested mixtures, while increasing the text-only dataset size provides little or no additional improvement.

Because the experiments do not evaluate higher ASR proportions or curricula that change the mixture during training, the authors explicitly leave unresolved whether these alternatives could improve performance. The open problem is therefore to assess multitask schedules beyond the fixed mixtures studied in the paper.

References

These findings leave open the possibility that higher ASR proportions---or perhaps annealing curricula that increase the ASR proportion towards the end of training---might yield stronger performance.

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models  (2609.09554 - Singh et al., 9 Sep 2026) in Appendix, Section “Multitask Learning Mixing Ratios” (Appendix: MTL Mixing Ratios)