ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition (2507.08477v1)

Published 11 Jul 2025 in cs.CL

Abstract: The deep integration of LLMs and automatic speech recognition systems has become a promising research direction with high practical value. To address the overfitting issue commonly observed in Low-Rank Adaptation (LoRA) during the supervised fine-tuning (SFT) stage, this work proposes an innovative training paradigm Iterative LoRA Training (ILT) in combination with an Iterative Pseudo Labeling strategy, effectively enhancing the theoretical upper bound of model performance. Based on Whisper-large-v3 and Qwen2-Audio, we conduct systematic experiments using a three-stage training process: Focus Training, Feed Back Training, and Fix Training. Experimental results demonstrate the effectiveness of the proposed method. Furthermore, the MegaAIS research team applied this technique in the Interspeech 2025 Multilingual Conversational Speech LLMing Challenge (MLC-SLM), achieving 4th in Track 1 (Multilingual ASR Task) and 1st place in Track 2 (Speech Separation and Recognition Task), showcasing the practical feasibility and strong application potential of our approach.

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Follow-up Questions

We haven't generated follow-up questions for this paper yet.

Generate Now

Authors (5)

YouTube

Show All Videos

ILT-Iterative LoRA Training through Focus-Feedback-Fix for Multilingual Speech Recognition (2507.08477v1)

Summary

Follow-up Questions

Related Papers

Authors (5)

YouTube

Don't miss out on important new AI/ML research