---
title: Speech-First Multimodal Training
url: https://www.emergentmind.com/topics/speech-first-multimodal-training-sfmt
type: topic
---

# Speech-First Multimodal Training

Speech-First Multimodal Training (SFMT) denotes a class of multimodal learning strategies in which speech is treated as the primary or anchoring modality, while auxiliary modalities such as text, video, phonology, articulatory measurements, or gesture are used to improve alignment, supervision, robustness, or downstream generalization. In the narrowest and most explicit sense, SFMT is a curriculum-learning method for Automated Speaking Assessment (ASA) that first optimizes an audio-only pathway and then introduces audio-text joint training [2508.12591]. Across adjacent literatures, the term also functions as a broader organizing principle for speech-anchored multimodal pre-training, cross-modal distillation, speech-guided fusion, modality expansion, and multimodal-training/unimodal-deployment schemes [2502.05766][2403.19822][2501.18157][2410.03798]. This broader usage suggests that SFMT is best understood not as a single architecture, but as a recurrent design pattern in which speech-derived representations constrain how other modalities are introduced, aligned, or discarded.

## 1. Terminology, scope, and historical lineage

The most explicit formulation of SFMT appears in ASA with Multimodal Large Language Models (MLLMs), where it is defined as a modality-level curriculum: Stage 1 uses audio only, and Stage 2 uses joint audio and text, with the aim of establishing “more robust modeling foundations of speech before cross-modal synergetic fusion” [2508.12591]. The motivating diagnosis is that naive multimodal training in MLLMs can become text-biased, thereby weakening delivery assessment, even though delivery depends strongly on acoustic and prosodic information.

Earlier and parallel work supplies the conceptual lineage for this speech-first perspective. A multimodal chain architecture from 2019 coupled ASR, TTS, image captioning, and image retrieval in a single framework, showing that ASR could be further trained without speech and text data and that cross-modal data augmentation could improve ASR performance [1906.00579]. In speech translation, the Fuse-Speech-Text model formalized speech, text, and fused speech-text translation within a single cross-modal system and explicitly analyzed modality gaps at the levels of input representation, semantics, and hidden states [2305.14042]. A survey of Speech Foundation Model (SFM) plus LLM systems for speech-to-text translation later organized the design space into five recurrent building blocks—SFM, length adapter, modality adapter, prompt-speech mixer, and LLM—while also emphasizing that fragmented evaluation makes it difficult to identify best-performing design choices [2402.12025].

Taken together, these strands suggest a broad encyclopedic definition: SFMT refers to multimodal training regimes in which speech is not merely one modality among many, but the modality around which training curricula, adapters, distillation targets, or inference paths are organized.

## 2. Canonical training paradigms

A central characteristic of SFMT is staged or asymmetric optimization. In ASA, SFMT is implemented as a two-stage curriculum over modalities. The first stage trains on audio only, replacing transcripts with null or junk values and updating the LoRA audio adapter to maximize acoustic-feature extraction; the second stage initializes from the first stage and fine-tunes on joint audio and text. The training objectives are written as
$$
\theta_{\text{LoRA}}^{1}
=
\arg\min_{\theta_{\text{LoRA}}}
\sum_{(\mathbf{a}, I, y)\in \mathcal{D}_{\text{audio}}}
\mathcal{L}\big(f_{\text{Phi-4}}(\mathbf{a}, I; \theta_{\text{LoRA}}), y\big)
$$
and
$$
\theta_{\text{LoRA}}^{2}
=
\arg\min_{\theta_{\text{LoRA}}^{1}}
\sum_{(\mathbf{a}, \mathbf{t}, I, y)\in \mathcal{D}_{\text{multi}}}
\mathcal{L}\big(f_{\text{Phi-4}}(\mathbf{a}, \mathbf{t}, I; \theta_{\text{LoRA}}^{1}), y\big),
$$
with the stated purpose of preserving acoustic discrimination while adding cross-modal integration [2508.12591].

A second paradigm is multi-stage pre-training followed by speech-only specialization. For ASR, one approach combines unsupervised audio-video pre-training with Masked Autoencoding (MAE), Contrastive Learning (CLR), or MAE+CLR, then performs supervised translation-based mid-training, and finally fine-tunes on downstream audio tasks. This framework reports relative word error rate improvements of up to 38.45% over baselines on LibriSpeech and SUPERB, and it explicitly treats audio as the retained modality after multimodal pre-training [2403.19822]. A third paradigm is self-powered modality expansion for large speech-text models: speech-text instruction tuning is augmented with model-generated instruction-following targets so that training shifts away from a pure ASR mapping and toward explicit conditioning on textual instructions, thereby mitigating “speech anchor bias” [2410.03798].

| Pattern | Representative formulation | Defining property |
|---|---|---|
| Modality curriculum | Audio only, then audio+text [2508.12591] | Speech foundation before fusion |
| Multi-stage pre-train/mid-train/fine-tune | Audio-video pre-training, translation mid-training, audio fine-tuning [2403.19822] | Auxiliary modalities enrich, speech remains deployment path |
| Self-powered instruction tuning | Model-generated instruction targets from ASR corpora [2410.03798] | Counteracts speech anchor bias |
| Synthetic parallel pre-training | Synthetic text, speech, and gesture corpus for pre-training [2404.19622] | Speech-centered multimodal scale-up |

These variants differ in architecture and task, but they share a structural asymmetry: speech is stabilized first, or speech-derived features serve as the supervisory target around which other modalities are organized.

## 3. Mechanisms for cross-modal transfer

SFMT systems use a limited set of recurring transfer mechanisms: contrastive alignment, representational distillation, cross-attention fusion, modality injection, and structured priors. In real-time MRI vocal tract segmentation, speech-guided multimodal learning is implemented as a three-stage framework. First, each frame is annotated with a 15-dimensional multi-hot phonological vector, which is converted into subject-specific four-channel spatial bounding-box priors for tongue, velum, upper lip, and lower lip. Second, a ViT-Base visual encoder and a WavLM-Base audio encoder are aligned through dual-level cross-modal contrastive pretraining, using both global-to-global and local-to-global alignment in a shared 256-dimensional space. Third, a six-layer cross-attention decoder lets visual tokens attend to audio tokens during training, but inference requires only the rtMRI image [2605.18466].

A closely related transfer logic appears in audiovisual representation learning via knowledge distillation from speech foundation models. Here, WavLM and iFLYTEK-speech act as teachers, and an audio-visual student is trained with a feature regression loss and a KL-divergence loss based on soft cluster labels. The teacher receives clean audio, whereas the student receives noisy audio and/or video; distillation is used both in pretraining and, when data are sufficient, during finetuning. A multi-teacher ensemble further improves robustness, and the reported performance is superior or at least comparable to previous state-of-the-art baselines across ASR, VSR, and AVSR [2502.05766].

Frozen speech backbones with lightweight multimodal injection constitute another major pattern. UASR-LLM adapts a frozen SFM to unified VSR, ASR, and AVSR by inserting visual injection modules into multiple SFM layers, using cross-attention with residual tanh-gated fusion. Training proceeds in two stages: visual injection pretraining by teacher-student distillation into the SFM latent space, followed by speech recognition finetuning through a feed-forward adaptor into a decoder-only LLM with LoRA updates [2510.22961]. FastSLM addresses the same alignment problem from the opposite end—sequence compression—by using Whisper-large-v3, a Hierarchical Frame Querying Transformer, and Qwen3-4B, compressing frame-level speech from 50 tokens per second to 1.67 tokens per second through hierarchical querying and three-stage training [2601.06199].

These mechanisms differ in whether they align modalities in feature space, inject one modality into another model’s hidden states, or distill multimodal knowledge into a speech-dominant path. The common denominator is that speech or speech-derived latent structure sets the geometry to which other modalities are mapped.

## 4. Multimodal training and reduced-modality inference

One of the strongest practical motifs in SFMT is that multimodal supervision during training need not imply multimodal requirements at deployment. In vocal tract segmentation, the rtMRI framework uses image, synchronized acoustic signal, and phonological descriptors at training time, but only rtMRI at inference. The paper states that image-only inference retains nearly all segmentation performance, with only a -0.22 Dice drop, while offering much lower computational latency than full multimodal input [2605.18466]. This is not framed as simple modality removal; rather, multimodal and phonological supervision are treated as knowledge that becomes internalized by the visual encoder and decoder.

MUTUD generalizes this principle under the name “Multimodal Training and Unimodal Deployment.” Its Temporally Aligned Modality feature Estimation (TAME) module estimates missing-modality features from the modality available at inference, using modality-specific codebooks and temporal alignment structure. Across audiovisual speech tasks, MUTUD is reported to reduce the performance gap between multimodal and corresponding unimodal models “to a considerable extent,” while reducing model size and compute compared to multimodal models, “in some cases by almost 80%” [2501.18157]. In speech enhancement at \(-5\) dB SNR, the reported STOI values are 80.1% for audio-only, 81.8% for MUTUD, and 83.3% for the full audiovisual model, with 3.6M parameters for MUTUD versus 15.7M for the audiovisual system [2501.18157].

A related but simpler form appears in multi-stage ASR pre-training, where video is used only during pre-training, after which the audio encoder is retained and fine-tuned for downstream speech tasks [2403.19822]. This suggests that SFMT frequently functions as a training-time asymmetry: auxiliary modalities are pedagogical rather than operational.

## 5. Task domains and empirical record

SFMT-related methods are now distributed across assessment, recognition, segmentation, synthesis, and speech-language modeling. In ASA, MLLM-based systems are reported to elevate holistic assessment performance from a PCC value of 0.783 to 0.846, while SFMT specifically improves the delivery aspect by an absolute 4% in accuracy over conventional training approaches. Delivery PCC rises to 0.848 and Macro Acc to 46.8% under the speech-first curriculum, compared with 0.831 and 42.3% for naive Phi-4 multimodal training [2508.12591]. In pronunciation assessment and mispronunciation detection, a LoRA fine-tuned Phi-4-multimodal-instruct model performs APA and MDD simultaneously, with predicted pronunciation scores showing PCC \(> 0.7\) with human-assigned scores and both WER and PER below 0.15; notably, fine-tuning only the LoRA layers is reported to be sufficient to reach performance comparable to updating all audio layers [2509.02915].

In articulatory and silent-speech-adjacent settings, speech-first transfer is used to compensate for limited articulatory data. A multimodal pre-training framework for articulatory-to-acoustic synthesis reports that utilizing the proposed transfer learning methods improves MRI-to-speech performance by 36% word error rate relative to prior work, and that multimodal pre-trained models consistently outperform unimodal baselines on three objective and subjective synthesis quality metrics [2412.13387]. The same study also reports that 33.1% of EMG spectrogram bins achieve Pearson \(r \geq 0.5\) when regressed from WavLM features, compared to 0% with a random control, which the authors interpret as evidence that latent information from acoustic models can be transferred to EMG representations [2412.13387].

In speech-and-gesture synthesis, synthetic data are used to create a large multimodal parallel corpus by chaining GPT-4 text generation, XTTS speech synthesis, Whisper ASR, forced alignment, and a diffusion-based gesture system. Pre-training on this synthetic data improves both the speech and motion synthesized by the multimodal model, and the proposed MAGI architecture further improves results when combined with synthetic pre-training. The reported WER for MAGI decreases from 13.28% to 9.29% after synthetic-data pre-training, while speech and gesture MOS also improve [2404.19622]. The paper explicitly describes this methodology as matching and extending a “speech-first multimodal” paradigm in which speech is the backbone for additional behaviors [2404.19622].

In audiovisual speech processing, the empirical record is similarly strong. The SFM-distillation approach reaches 37.2 VSR WER, 3.0 ASR WER, and 2.8 AVSR WER on LRS3 with 30 hours of labeled data, outperforming audiovisual self-supervised baselines in that setting [2502.05766]. UASR-LLM reports best WERs of 20.9% for VSR, 0.84% for ASR, and 0.69% for AVSR on LRS3 test under its largest setting, while also maintaining strong robustness under noisy conditions [2510.22961]. FastSLM, oriented toward long-form speech understanding and reasoning, reports representation of speech with only 1.67 tokens per second while remaining competitive on ASR, automatic speech translation, spoken summarization, and spoken-query question answering [2601.06199].

## 6. Methodological tensions, misconceptions, and open questions

A common misconception is that “speech-first” means speech-only. The literature does not support that reading. SFMT systems often depend on text, video, phonology, gesture, or articulatory measurements during training; the defining feature is not modal exclusivity but modal hierarchy. Speech is the anchor around which other modalities are aligned, injected, or distilled, and in several systems the auxiliary modalities disappear at inference time [2605.18466][2501.18157][2403.19822].

Another misconception is that multimodal fusion always fails in the same way. The literature instead shows two opposite pathologies. In ASA, naive multimodal training is described as text-biased, with the model gravitating toward text and underutilizing acoustic information needed for delivery assessment [2508.12591]. In large speech-text models, the converse failure appears as “speech anchor bias,” where the model over-relies on the speech input, interprets the entire speech modality as directive, and neglects textual instructions [2410.03798]. This opposition is important: SFMT is not simply about privileging speech more strongly, but about structuring training so that speech remains informative without suppressing legitimate cross-modal conditioning.

Open questions remain substantial. The survey of SFM+LLM speech translation concludes that no consensus exists on the best SFM, length adapter, modality adapter, prompt-speech mixer, or LLM; public comparisons are hindered by inconsistent datasets, metrics, and language pairs, and the survey recommends standardized public setups, broader use of semantic metrics such as COMET alongside BLEU, controlled architectural ablations, and stronger analysis of multilinguality, efficiency, and in-context learning [2402.12025]. A plausible implication is that SFMT has matured into a recognizable paradigm before it has converged into a standardized methodology. Its future development is therefore likely to depend less on introducing new modalities than on clarifying when speech-first curricula, distillation routes, and deployment asymmetries are preferable to symmetric multimodal training.

Source: https://www.emergentmind.com/topics/speech-first-multimodal-training-sfmt