---
title: 'HM-Talker: Speech Recog & Talking-Head Synthesis'
url: https://www.emergentmind.com/topics/hm-talker
type: topic
---

# HM-Talker: Speech Recog & Talking-Head Synthesis

Searching arXiv for HM-Talker and closely related papers to ground the article.
arxiv_search.query({"search_query":"all:\"HM-Talker\" OR ti:\"HM-Talker\" OR abs:\"speaker mask transformer\" OR ti:\"Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions\"","start":0,"max_results":10,"sort_by":"submittedDate","sort_order":"descending"})
I found the relevant arXiv entries for HM-Talker and adjacent multi-talker research, including the talking-head synthesis paper "HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis" [2508.10566], the multi-talker speech-recognition paper summarized as HM-Talker / speaker-mask transformer [2312.10959], and related multi-talker ASR work such as MT-LLM [2409.08596], Whisper-based multilingual multi-talker transfer [2305.18747], and joint multi-/target-talker Whisper adaptation [2407.09817].
HM-Talker is a name that has been used in arXiv literature for two distinct research systems in different subfields. In speech processing, it denotes a transformer-based method for multi-talker overlapped speech recognition that combines speaker-aware autoregressive transcription with a dedicated speaker mask branch for diarization-oriented segment detection [2312.10959]. In talking-face generation, it denotes "HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis," an audio-driven talking head framework that combines implicit audio motion cues with explicit Action Unit (AU) priors through cross-modal disentanglement and hybrid fusion [2508.10566]. The shared label therefore does not identify a single canonical architecture; rather, it refers to separate systems concerned with speech-related human communication signals in either acoustic or visual form.

## 1. Nomenclature and scope

The speech-recognition usage of HM-Talker arises from the paper "Speaker Mask Transformer for Multi-talker Overlapped Speech Recognition" [2312.10959]. In that work, HM-Talker refers to the proposed speaker-mask transformer method for multi-talker overlapped speech recognition. Its stated objective is not only to recognize the lexical content of overlapped speech, but also to address speaker diarization by answering **WHAT** each speaker said, **WHO** said it, and **WHEN** each speaker was active [2312.10959].

A separate usage appears in the later paper "HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis" [2508.10566]. There, HM-Talker is an audio-driven talking head synthesis framework that targets motion blur, lip jitter, and phoneme-viseme misalignment by introducing a hybrid motion representation composed of implicit and explicit motion cues [2508.10566].

A common misconception is to treat HM-Talker as a single model family. The arXiv record instead supports a narrower and more precise view: the name has been independently attached to two technically unrelated systems, one for overlapped speech recognition and diarization, and one for talking head video generation [2312.10959] [2508.10566].

## 2. HM-Talker in multi-talker overlapped speech recognition

In the speech-recognition setting, HM-Talker is formulated as a joint ASR and speaker diarization system built on **Whisper base**, with **74M parameters**, **6 transformer blocks in encoder and decoder**, **hidden size: 512**, and **80-channel log-magnitude Mel spectrograms** using **25 ms** windows and **10 ms** stride [2312.10959]. The method modifies autoregressive transformer ASR by adding **speaker labels** to the serialized output and by attaching a **speaker mask branch** to the encoder representation.

The speaker-labeled output is introduced through token sequences such as
$$
\mathbf{U}_{\text{SPK-TS-1} = <S_1> <T^1_s> \mathbf{W}_1 <T^1_e> \dots \ <S_k> <T^k_s> \mathbf{W}_k <T^k_e> \dots
$$
and
$$
\mathbf{U}_{\text{SPK-TS-2} = <S_1> <T^1_s> <T^1_e> \mathbf{W}_1 \dots \ <S_k> <T^k_s> <T^k_e> \mathbf{W}_k \dots .
$$
The paper quantizes timestamps with **20 ms resolution**, following Whisper-style timestamp tokens [2312.10959]. Its ASR loss is written as
$$
L_{\text{ASR} = -\log P(\mathbf{U}  | \mathbf{X} , \mathbf{\Theta} )
$$
with \(\mathbf{X}\) the acoustic feature sequence and \(\mathbf{\Theta}\) the model parameters [2312.10959].

The central HM-Talker innovation is the **speaker mask branch**, inspired by **Mask R-CNN**. The ASR branch still predicts serialized text with speaker labels, but diarization is shifted to a separate branch that predicts the **speech activity mask** of each speaker from encoder hidden states. For speaker \(k\), the mask loss is
$$
L^k_{\text{mask} = -\frac{1}{D}\sum_{i=1}^{D} \left[ y_i \log(p_i) + (1 - y_i) \log(1 - p_i) \right]
$$
where \(y_i \in \{0,1\}\) is the ground-truth mask label and \(p_i\) is the sigmoid output [2312.10959]. These mask targets are constructed using an **energy-based VAD** over **20 ms segments**, with time-alignment timestamps from the **Montreal Forced Aligner** used to anchor speech starts [2312.10959].

The joint training objective is
$$
L_{\text{SPK-MASK} = (1-\lambda) * L_{ASR} + \lambda \sum_{k=1}^{K} L^k_{mask}
$$
and, for the proposed model, the decoder target removes timestamps and keeps speaker labels:
$$
\mathbf{U}_{\text{SPK} = <S_1> \mathbf{W}_1 \dots <S_k> \mathbf{W}_k \dots .
$$
The paper states that the mask branch is optimized **only when the main branch outputs speaker labels**, tying diarization learning directly to speaker-aware transcription [2312.10959].

This design is motivated by a specific failure mode of timestamp-token diarization. If one speaker is interrupted and later resumes, start/end timestamps alone can represent the serialized utterances awkwardly. The speaker mask branch instead learns a speaker-specific activity map from encoder features, which can indicate all speech frames belonging to a speaker even when those frames are non-contiguous in the output sequence [2312.10959]. This suggests that HM-Talker treats transcription and segmentation as related but not identical prediction problems.

## 3. Architectural variants, data, and empirical behavior of the speech-recognition system

The paper evaluates several mask-head configurations: **L-FC**, **L-FC-CNN**, **CA-FC**, and **CA-FC-CNN**. The CNN stack uses **32 and 64 kernels**, **kernel size \(2 \times 1\)**, **stride 1**, and **dropout 0.25 after the second CNN** [2312.10959]. The reported finding is that the **cross-attention + CNN** designs generally work best for diarization [2312.10959].

Training uses all **960 hours** of LibriSpeech training data and constructs two overlap regimes. **Case 1** is simple two-speaker overlap created by randomly selecting an utterance from a different speaker and mixing at **0 dB SIR**. **Case 2** is a more complex overlap pattern with utterance order **Speaker 1, Speaker 2, Speaker 1**, and overlap duration randomly sampled from **0 to 5 seconds**. The resulting training sets are **train960-Org**, **train960-Set1**, and **train960-Set2** with ratios **1:1** and **1:1:1** as specified in the data [2312.10959]. Evaluation is performed on **Set1-1s**, **Set1-3s**, and **Set2-1s** derived from **test-clean**, with diarization measured by **DER** with **0.2 collar**, speaker counting by **SCA**, and recognition by **WER** [2312.10959].

Representative results show the trade-off between timestamp-token methods and the mask branch. Under **Case 1**, **SPK** gives **Set1-1s WER: 5.93, SCA: 99.50** and **Set1-3s WER: 9.47, SCA: 99.09**. Timestamp-based **SPK-TS-1** reduces diarization error but worsens WER, with **Set1-1s WER: 8.14, DER: 2.72, SCA: 99.89** and **Set1-3s WER: 13.13, DER: 3.09, SCA: 99.90**. By contrast, **Mask: CA-FC-CNN (0.5)** achieves **Set1-1s WER: 6.03, DER: 1.54, SCA: 99.69** and **Set1-3s WER: 9.91, DER: 2.04, SCA: 99.65** [2312.10959].

The strongest contrast appears in **Case 2**, where the paper notes that **SPK-TS-2 performs badly in the hardest condition**, with **Set2-1s DER: 26.37%**. The proposed mask method remains low in DER: **Mask: L-FC (0.5)** gives **Set2-1s WER: 6.59, DER: 1.10, SCA: 99.95**, while **CA-FC-CNN (0.5)** gives **Set2-1s WER: 6.51, DER: 0.70, SCA: 99.85** [2312.10959]. With **Whisper Large-v2**, the paper reports that ASR WER improves significantly while DER stays low; for example, **LargeV2, Mask: L-FC (0.5)** gives **Set1-1s DER: 0.76** and **Set2-1s DER: 0.80** [2312.10959].

The paper also studies the loss balance \(\lambda\). It reports that **small \(\lambda\)** gives better ASR and weaker diarization, **large \(\lambda\)** gives stronger diarization and more ASR degradation, and **\(\lambda = 0.5\)** gives the best trade-off [2312.10959]. This supports the interpretation that diarization supervision is helpful but can interfere with lexical modeling if weighted too heavily.

Within the broader multi-talker ASR literature, HM-Talker occupies a distinct position. Whisper adaptation with **enhanced Serialized Output Training (SOT)** and timestamps has been used to jointly model multi-talker ASR, speaker counting, and utterance timestamp prediction in meeting transcription [2305.18747]. Whisper has also been extended with a **Sidecar Separator**, **Target Talker Identifier**, and **soft prompt tuning** to jointly address multi-talker and target-talker ASR [2407.09817]. MT-LLM reformulates overlapped speech transcription as instruction-conditioned language generation, supporting multi-talker ASR, target-talker ASR, sex-specific ASR, order-specific ASR, target-lingual ASR, and keyword-tracing ASR [2409.08596]. Relative to those directions, HM-Talker is specifically focused on the claim that timestamp tokens alone are brittle for diarization in complex overlap and that explicit mask prediction offers a better solution [2312.10959].

## 4. HM-Talker in audio-driven talking head synthesis

The later HM-Talker is an audio-driven talking head generation framework targeting **motion blur**, **lip jitter**, and **phoneme-viseme misalignment** [2508.10566]. Its central proposal is a **hybrid motion representation** that combines **implicit motion cues** with **explicit motion cues** based on **Action Units (AUs)**. The paper states that explicit cues use anatomically defined facial muscle movements alongside implicit features to minimize phoneme-viseme misalignment [2508.10566].

The architecture has two named components. The **Cross-Modal Disentanglement Module (CMDM)** disentangles implicit and explicit motion information across audio and video modalities. CMDM takes portrait video frames \(\mathcal{I}_{1:T}\), audio features from an audio-visual encoder inherited from **SyncTalk**, OpenFace-extracted AUs, and an MLP-based audio-to-AU projection branch [2508.10566]. It outputs four motion features:

- \(\mathbf{c}_{v,u}^{e} \in \mathbb{R}^7\): upper-face explicit feature.
- \(\mathbf{c}_{v,l}^{e} \in \mathbb{R}^{32}\): lower-face explicit feature from video.
- \(\mathbf{c}_{a,l}^{i} \in \mathbb{R}^{32}\): lower-face implicit audio feature.
- \(\mathbf{c}_{a,l}^{e} \in \mathbb{R}^{32}\): audio-predicted explicit feature [2508.10566].

The explicit AU representation uses **17 AUs** split into upper-face and lower-face groups:
$$
\mathcal{A}_u = \{AU_{01}, AU_{02}, AU_{04}, AU_{05}, AU_{06}, AU_{07}, AU_{45}\}
$$
and
$$
\mathcal{A}_l = \{AU_{09}, AU_{10}, AU_{12}, AU_{14}, AU_{15}, AU_{17}, AU_{20}, AU_{23}, AU_{25}, AU_{26}\}.
$$
Upper-face explicit features are formed as
$$
\mathbf{c}_{v,u}^{e} = \bigoplus_{i \in \mathcal{A}_u} AU_i \in \mathbb{R}^7,
$$
while lower-face explicit features are
$$
\mathbf{c}_{v,l}^{e} = \text{MLP}(\mathcal{A}_l) \oplus \mathcal{A}_l \in \mathbb{R}^{32}.
$$
The residual connection is used to preserve raw AU semantics while allowing nonlinear AU interactions [2508.10566].

The implicit audio feature is produced through **AudioNet + AudioAttNet** from a pre-trained audio-visual encoder feature \(a \in \mathbb{R}^{512}\), yielding \(\mathbf{c}_{a,l}^{i} \in \mathbb{R}^{32}\) [2508.10566]. CMDM then maps audio-only features into AU-like explicit space through the Audio-to-AU Mapper:
$$
\mathbf{c}_{a,l}^{e} = \sigma_2(\mathbf{W}_2(\sigma_1(\mathbf{W}_1\mathbf{c}_{a,l}^{i} + \mathbf{b}_1))+\mathbf{b}_2),
$$
with cross-modal alignment enforced by
$$
\mathcal{L}_{align} = \mathcal{L}_1(\mathbf{c}_{a,l}^{e}, \mathbf{c}_{v,l}^{e}) .
$$
This makes explicit articulatory priors available at inference even when lower-face motion is driven only by audio [2508.10566].

The second major component is the **Hybrid Motion Modeling Module (HMMM)**, which fuses implicit and explicit motion cues through gated fusion:
$$
\mathbf{c}_f = \mathcal{G}(\mathbf{c}_{a,l}^{i*},\mathbf{c}_{\cdot,l}^{e})=\alpha \odot \mathbf{c}_{\cdot,l}^{e} + (1 - \alpha) \odot \mathbf{c}_{a,l}^{i*},
$$
with
$$
\alpha = \text{MLP}_g(\mathbf{c}_{a,l}^{i*}\oplus \mathbf{c}_{\cdot,l}^{e}) .
$$
During training, HMMM randomly selects one of three fusion paths—**audio path**, **masked path**, and **vanilla path**—with default ratio
$$
\mathcal{P}_{audio} : \mathcal{P}_{masked} : \mathcal{P}_{vanilla} = 4:4:2 .
$$
The paper states that this random pairing strategy is designed to mitigate identity-dependent biases in explicit features and enforce identity-agnostic learning [2508.10566].

HM-Talker extends **TalkingGaussian**, a **3D Gaussian Splatting-based** talking head system. Lower-face and upper-face features are modulated by region-specific attention from Gaussian positional encoding \(\mathcal{H}(\mu)\):
$$
\mathbf{C}_f = \mathbf{c}_f \odot \text{MLP}_f(\mathcal{H}(\mu)), \qquad
\mathbf{C}_u = \mathbf{c}_{v,u}^{e} \odot \text{MLP}_u(\mathcal{H}(\mu)),
$$
followed by deformation prediction
$$
\delta_{face} = \text{MLP}(\mathcal{H}(\mu) \oplus \mathbf{C}_u \oplus \mathbf{C}_f).
$$
Final blending is written as
$$
\hat{I}_{head} = C_{face} \times A_{face} + C_{mouth} \times (1 - A_{mouth}) .
$$
Training follows **three-stage optimization**—**static initialization**, **motion learning**, and **fine-tuning**—with the face branch and inside-mouth branch trained in parallel for **50,000 iterations** and then jointly fine-tuned for **15,000 iterations**, using **Adam and AdamW**, **learning rate \(5 \times 10^{-4}\)**, and loss weights
$$
\lambda_1 = 0.2,\quad \lambda_2 = 0.5,\quad \lambda_3 = 10^{-3}.
$$
The total objective is
$$
\mathcal{L}_{DF} = \mathcal{L}_{1} + \lambda_1 \mathcal{L}_{D\text{-}SSIM} + \lambda_2 \mathcal{L}_{LPIPS} + \lambda_3 \mathcal{L}_{align} .
$$
Although training uses both audio and image inputs, evaluation and deployment are **audio-only for lower-face motion** [2508.10566].

## 5. Evaluation profile of the talking-head system

The talking-head HM-Talker is evaluated on **five public video sequences**: **Lieu**, **Jae-in**, **Obama**, **May**, and **Shaheen**. The data contain **3 male and 2 female subjects**, have average duration **7,637 frames**, run at **25 fps**, and are mostly **512×512** resolution, with some **450×450** videos. A **10:1 train-validation split** is used [2508.10566].

The comparison set spans **2D methods** (**IP-LAP**, **TalkLip**, **DINet**), **NeRF-based methods** (**AD-NeRF**, **RAD-NeRF**, **ER-NeRF**, **SyncTalk**), and **3DGS-based methods** (**GaussianTalker**, **TalkingGaussian**) [2508.10566]. The evaluation metrics are divided into rendering quality (**PSNR**, **SSIM**, **LPIPS**), motion quality (**LMD**, **AUE-(L/U)**, **Sync-C**), lip-sync experiments with **Sync-D**, and efficiency in terms of **training time** and **FPS** [2508.10566].

For self-reconstruction, HM-Talker reports **PSNR: 35.15**, **LPIPS: 0.0207**, **SSIM: 0.9971**, **LMD: 2.514**, **AUE-L / AUE-U: 0.53 / 0.22**, **Sync-C: 7.807**, **Training time: 0.51h**, and **FPS: 110** [2508.10566]. The paper compares this with **SyncTalk: 34.51 PSNR, 0.0221 LPIPS, 7.502 Sync-C** and **TalkingGaussian: 32.48 PSNR, 0.0309 LPIPS, 6.246 Sync-C** [2508.10566]. For unseen audio speakers, the reported lip-sync results are **Shaheen audio: Sync-D 7.590, Sync-C 7.972** and **Lieu audio: Sync-D 7.292, Sync-C 7.994** [2508.10566].

The ablations are used to justify the module design. Replacing CMDM with **ExpNet** from **SadTalker** causes performance to drop, especially in lip synchronization. Replacing AUs with **3DMM** or **BlendShape** works reasonably, but **AU performs best for articulation control**. Replacing the inherited audio-visual encoder with **DeepSpeech** or **HuBERT** shows **DeepSpeech performs worst** and **HuBERT is better but still below AVE**. For HMMM, **gated fusion is best overall**, while **purely explicit**, **purely implicit**, and **fixed alpha** are weaker. The default three-path stochastic training ratio **\(4:4:2\)** is also reported as best, with masking rate sampled uniformly from **0.1 to 0.3** and best masking rate around **0.2** [2508.10566].

The paper does not provide an extensive limitations section, but it explicitly frames several constraints through method description and experiment scope: dependence on training-time visual supervision, subject-specific portrait videos, reliance on **OpenFace** AU extraction, emphasis on lower-face articulation, and a secondary role for the inside-mouth branch [2508.10566]. A plausible implication is that the method is strongest as a subject-specific, AU-informed audio-driven avatar system rather than as a fully open-domain talking-face generator.

Within the broader talking-head literature, HM-Talker can be read alongside diffusion-based motion-structured systems such as **MoDiTalker**, which separates talking-head generation into **Audio-to-Motion (AToM)** and **Motion-to-Video (MToV)** and emphasizes residual landmark diffusion, disentangled attention, and tri-plane-conditioned video synthesis [2403.19144]. The two papers belong to different model classes—3DGS-based hybrid motion fusion in HM-Talker versus diffusion-based motion-video disentanglement in MoDiTalker—but both treat explicit intermediate motion structure as preferable to direct audio-to-RGB mapping [2508.10566] [2403.19144].

## 6. Conceptual comparison and significance

The two HM-Talker systems address different modalities, yet they share a methodological preference for **explicit structure in addition to sequence generation**. In the speech-recognition HM-Talker, speaker-aware autoregressive decoding is supplemented by a **speaker mask branch** because timestamp tokens alone are described as insufficient for robust diarization in complex overlap [2312.10959]. In the talking-head HM-Talker, implicit audio features are supplemented by **explicit AUs** because purely implicit audio-facial mapping is described as a source of motion blur, lip jitter, and phoneme-viseme misalignment [2508.10566].

This suggests a common design intuition: latent sequence models are useful for content generation, but difficult subproblems such as diarization or articulatory control benefit from dedicated explicit representations. In one case the explicit representation is a binary speaker activity mask; in the other it is an AU-based anatomical prior [2312.10959] [2508.10566].

The two usages also occupy different positions relative to adjacent research. The speech HM-Talker is part of a line that includes SOT-based Whisper adaptation for multilingual meeting transcription [2305.18747], Sidecar- and TTI-based Whisper adaptation for joint multi-talker and target-talker ASR [2407.09817], and instruction-following speech-language modeling for cocktail-party transcription in MT-LLM [2409.08596]. The talking-head HM-Talker belongs to a different trajectory that includes hybrid explicit/implicit motion modeling, 3D Gaussian Splatting renderers, and cross-modal motion control, with conceptual overlap with motion-disentangled video generation methods such as MoDiTalker [2508.10566] [2403.19144].

For encyclopedia purposes, the most precise interpretation is therefore terminological rather than genealogical: **HM-Talker is a reused research name, not a unified benchmark or single algorithmic lineage**. In speech literature it denotes a joint ASR-plus-diarization transformer centered on speaker masks [2312.10959]. In talking-head synthesis it denotes a hybrid-motion, AU-guided 3DGS framework for high-fidelity lip-synchronized video generation [2508.10566].

Source: https://www.emergentmind.com/topics/hm-talker