Papers
Topics
Authors
Recent
Search
2000 character limit reached

HuBERT: Self-Supervised Speech Recognition Model

Updated 15 September 2026
  • HuBERT is a self-supervised speech model that combines hidden-unit prediction with masked language modeling paradigms to produce powerful speech representations.
  • HuBERT achieves state-of-the-art performance in automatic speech recognition (ASR) and other speech tasks, outperforming wav2vec 2.0 in several settings and demonstrating robustness in low-resource and multi-task environments.
  • HuBERT iteratively refines its hidden-unit predictions through a process of offline clustering and self-supervision, achieving high temporal consistency crucial for learning complex speech patterns.
  • The model's complex masking strategy and iterative clustering improve not only ASR but also auxiliary speech tasks such as speaker verification and emotion recognition, making it highly versatile.

HuBERT (Hidden-Unit BERT) is a self-supervised speech representation-learning model that trains a convolutional waveform encoder and Transformer context network to predict discrete hidden-unit labels at masked temporal positions. The labels are generated by offline clustering rather than supplied by phonetic or textual annotation. HuBERT’s defining methodology combines temporally aligned pseudo-labels, span masking, masked-only prediction, and iterative reclustering of learned representations. The model was introduced for automatic speech recognition (ASR), where it matched or improved upon wav2vec 2.0 on several LibriSpeech and Libri-light settings, including low-resource fine-tuning (Hsu et al., 2021). The name is also used for a separate monolingual Hungarian BERT model, huBERT, whose architecture and purpose are unrelated to the speech model (Ács et al., 2021).

1. Terminology and historical context

The speech model is conventionally written HuBERT, whereas huBERT denotes a Hungarian contextualized LLM. HuBERT adapts the masked-language-modeling paradigm of BERT to continuous speech: the input is a waveform-derived sequence, the targets are automatically clustered acoustic units, and prediction is performed at frame positions rather than word-piece positions (Hsu et al., 2021).

HuBERT was proposed in the context of self-supervised speech representation learning, alongside systems such as wav2vec 2.0 and data2vec. Its principal distinction from wav2vec 2.0 is the separation of unit discovery and representation learning. wav2vec 2.0 uses learned vector quantization, contrastive prediction, negative samples, and a diversity loss; HuBERT uses offline kk-means clustering, masked cross-entropy, and iterative target refinement. In abstract form, HuBERT implements:

unit discovery⟶masked contextual prediction.\text{unit discovery} \longrightarrow \text{masked contextual prediction}.

The initial HuBERT Base model has 12 Transformer layers, hidden size 768, and approximately 95 million parameters. HuBERT Large has 24 Transformer layers, hidden size 1024, and approximately 317 million parameters; HuBERT X-Large has 48 Transformer layers, hidden size 1280, and approximately 964 million parameters (Hsu et al., 2021). The principal pretraining corpora are LibriSpeech, containing 960 hours, and Libri-light, containing approximately 60,000 hours.

The model’s pretraining targets are not guaranteed to be phonemes, words, or linguistically meaningful categories. They are discrete cluster assignments derived from acoustic or learned representations. Their usefulness depends primarily on temporal consistency and their suitability for masked prediction rather than on initial phonetic correctness.

2. Architecture and masked hidden-unit prediction

HuBERT processes raw waveform audio with a seven-layer one-dimensional convolutional feature encoder. The layers use 512 channels, strides

[5,2,2,2,2,2,2],[5,2,2,2,2,2,2],

and kernel widths

[10,3,3,3,3,2,2].[10,3,3,3,3,2,2].

The total downsampling factor is $320$, producing one latent vector approximately every 20 ms for 16-kHz audio. The convolutional output is passed to a bidirectional Transformer context network.

Let the waveform-derived representation be

C(X)=[c1,…,cT],C(X)=[c_1,\ldots,c_T],

and let MM be the set of masked frame positions. HuBERT replaces masked representations with a learned mask embedding, producing X~\widetilde{X}. The Transformer generates contextual states oto_t, from which the model predicts the cluster identity ztz_t:

unit discovery⟶masked contextual prediction.\text{unit discovery} \longrightarrow \text{masked contextual prediction}.0

where unit discovery⟶masked contextual prediction.\text{unit discovery} \longrightarrow \text{masked contextual prediction}.1 is a learned projection, unit discovery⟶masked contextual prediction.\text{unit discovery} \longrightarrow \text{masked contextual prediction}.2 is a cluster embedding, unit discovery⟶masked contextual prediction.\text{unit discovery} \longrightarrow \text{masked contextual prediction}.3 is cosine similarity, and unit discovery⟶masked contextual prediction.\text{unit discovery} \longrightarrow \text{masked contextual prediction}.4.

The masked-region log-likelihood is

unit discovery⟶masked contextual prediction.\text{unit discovery} \longrightarrow \text{masked contextual prediction}.5

Training is ordinarily expressed as minimizing the negative log-likelihood:

unit discovery⟶masked contextual prediction.\text{unit discovery} \longrightarrow \text{masked contextual prediction}.6

HuBERT’s standard masking configuration selects approximately 8% of frame positions as span starts and masks 10 consecutive frames from each selected start. At a 20-ms frame rate, a span covers approximately 200 ms, although overlap and boundary effects alter the exact duration (Hsu et al., 2021).

Masked-only prediction is central to the method. If the loss were applied primarily to unmasked frames, the model could reproduce the local acoustic evidence or imitate errors in the clustering teacher. With masked-only prediction, it must infer the hidden unit from surrounding acoustic context, learning both local acoustic regularities and longer-range dependencies.

3. Offline clustering and iterative target refinement

The first HuBERT iteration uses 39-dimensional MFCC features, consisting of 13 MFCC coefficients, first-order derivatives, and second-order derivatives. A unit discovery⟶masked contextual prediction.\text{unit discovery} \longrightarrow \text{masked contextual prediction}.7-means model with 100 clusters assigns each frame a pseudo-label:

unit discovery⟶masked contextual prediction.\text{unit discovery} \longrightarrow \text{masked contextual prediction}.8

These initial labels are noisy and are not treated as ground-truth phoneme annotations.

After training the first HuBERT model, intermediate Transformer representations are extracted and reclustered. For the Base configuration, the second iteration uses 500 clusters derived from the sixth Transformer layer of the first model. The new cluster assignments supervise a subsequent HuBERT model. For Large and X-Large models trained on Libri-light, features from the ninth Transformer layer of the second-iteration Base model are used to generate labels (Hsu et al., 2021).

The iterative procedure can therefore be expressed as:

  1. Extract MFCC features and cluster them.
  2. Train a masked predictor on the resulting frame labels.
  3. Extract intermediate Transformer representations.
  4. Recluster those representations.
  5. Train a subsequent HuBERT model using the refined labels.

A central HuBERT claim is that cluster consistency is more important than initial cluster correctness. A noisy but stable target sequence provides a predictable learning problem. Once the model learns contextual structure, its hidden representations can support more informative clustering. Reported phone-normalized mutual information increases from approximately 0.253 for 100-cluster MFCC targets to approximately 0.684 for 500-cluster first-iteration HuBERT features (Hsu et al., 2021).

The number of clusters is not universally optimal. Larger codebooks can encode finer phonetic distinctions, but excessive granularity can capture speaker or within-phone variation. Cluster purity can also increase artificially with the number of clusters, so phone-normalized mutual information and downstream masked-prediction performance are more informative than purity alone.

The offline pipeline has computational consequences. Each refinement iteration requires feature extraction over a large corpus and another unit discovery⟶masked contextual prediction.\text{unit discovery} \longrightarrow \text{masked contextual prediction}.9-means stage. Academic-compute implementations reduce the burden through distributed and streaming [5,2,2,2,2,2,2],[5,2,2,2,2,2,2],0-means, feature sampling, Kaldiio-based storage, bfloat16 training, gradient accumulation, and dynamic batch construction. An ESPnet reproduction trained HuBERT Base with eight A100 GPUs rather than the 32 GPUs used in an original-style Base configuration, while obtaining an overall SUPERB score of 80.7, matching the original HuBERT Base result (Chen et al., 2023).

4. Fine-tuning and empirical performance

For ASR fine-tuning, HuBERT removes the pretraining projection layer and replaces it with a randomly initialized CTC softmax layer. The convolutional feature encoder is frozen in the reported recipe, while the Transformer is fine-tuned. The output vocabulary contains 26 English letters, space, apostrophe, and the CTC blank. Decoding may combine CTC probabilities, a LLM, and a word-length penalty (Hsu et al., 2021).

HuBERT Large reaches 7.6% test-other WER with 10 minutes of labeled LibriSpeech data, while HuBERT X-Large reaches 6.8% in the same setting. With the full 960 hours of labeled data, HuBERT X-Large reaches 2.9% test-other WER. In the reported comparisons, HuBERT matches or improves upon wav2vec 2.0 across most fine-tuning regimes. It is substantially better than DiscreteBERT in the 10-minute condition, where DiscreteBERT obtains 25.2% test-other WER and HuBERT Large obtains 7.6% (Hsu et al., 2021).

HuBERT representations are not uniformly most informative at the final Transformer layer. Morphological probing, POS tagging, and NER experiments on Hungarian demonstrate the broader methodological importance of layer selection: middle Transformer layers frequently provide the best linguistic representations, while the final layer is not necessarily optimal (Ács et al., 2021). In speech, later work similarly reports that HuBERT layer choice affects phonetic, syllabic, lexical, semantic, and paralinguistic behavior.

HuBERT transfers beyond ASR when fine-tuned. A benchmark on speech emotion recognition, speaker verification, and spoken language understanding reports the following best HuBERT results: 79.58% weighted accuracy for speaker-dependent emotion recognition, 73.01% for speaker-independent emotion recognition, 2.36% equal error rate on VoxCeleb1 speaker verification, 89.38% intent accuracy, and 78.92% slot F1 on SLURP (Wang et al., 2021). The optimal adaptation strategy is task-dependent: partial fine-tuning is favored for small emotion datasets, entire fine-tuning is favored for speaker verification, and the two strategies are similar for spoken language understanding.

HuBERT has also been used as a frozen feature extractor for acoustic landmark detection. Frozen HuBERT-base representations achieve F1@20 ms of 0.77 and F1@30 ms of 0.84 on a corpus with eight landmark types, outperforming the reported mel and wav2vec 2.0 configurations (Cámara et al., 22 Jun 2026). The strongest results occur for abrupt stop and fricative events, whereas vowels remain more difficult.

5. Representation granularity and model extensions

HuBERT’s conventional 20-ms temporal resolution is not universally optimal. Multi-resolution systems train HuBERT streams at 20, 40, and 100 ms and fuse them either in parallel or hierarchically. HuBERT-MR-P performs weighted fusion after temporal upsampling, whereas HuBERT-MR-H combines low- and high-resolution streams through convolutional and transposed-convolutional modules (Shi et al., 2023). The multi-resolution model improves several SUPERB tasks and LibriSpeech ASR results, although its parameter count is higher than that of a single Base model.

A related model, MR-HuBERT, integrates 20-ms and 40-ms streams during pretraining through a hierarchical Transformer and applies masked-unit prediction at both resolutions. Its default two-resolution configuration reduces reported MACs from 431G for HuBERT Base to 394G for the comparable MR-HuBERT Base, while the large configuration reduces MACs from 1116G to 971G (Shi et al., 2023). In 100-hour LibriSpeech ASR, MR-HuBERT-H with 100/40/20-ms streams reaches 6.11% WER without language-model rescoring and 3.31% with rescoring, compared with 7.73% and 3.81% for ordinary HuBERT Base.

Other extensions modify the learning objective rather than only the temporal architecture. MS-HuBERT addresses the mismatch between masked pretraining and unmasked inference by processing masked and complete views jointly and swapping their hidden representations at masked positions after each Transformer layer. It also applies masked prediction at multiple Transformer layers and cluster resolutions, using cluster inventories of 1000, 500, 250, 125, 50, and 25 units (Yadav et al., 2024). MS-HuBERT improves over vanilla HuBERT on the reported 10-hour and 100-hour ASR conditions and reaches 2.4% test-clean and 5.5% test-other WER with 960 hours of labeled data.

HuBERTopic adds a global utterance-level objective. It removes adjacent duplicate HuBERT units, treats the resulting sequences as pseudo-text, applies Latent Dirichlet Allocation, and assigns each utterance a topic label. A fixed random 512-dimensional CLS vector and an auxiliary topic-classification loss are added to HuBERT, with topic-loss weight [5,2,2,2,2,2,2],[5,2,2,2,2,2,2],1 (Maekaku et al., 2023). The method improves HuBERT on all eight reported SUPERB tasks in the LS-100h setting, although the discovered topic labels also correlate strongly with gender, speaker, book, and chapter. Consequently, “topic” denotes statistically consistent utterance-level information rather than necessarily human-interpretable semantic themes.

SD-HuBERT changes the granularity of supervision in the opposite direction from HuBERTopic. It fine-tunes pretrained HuBERT with sentence-level self-distillation, using an aggregator token and an exponential-moving-average teacher. The resulting representations exhibit syllable-like similarity blocks and low-norm boundary frames (Cho et al., 2023). A subsequent speaker-disentangled method removes the CLS token, introduces speaker perturbation, and trains with a frame-level BYOL-style objective. It obtains F1 70.3 for syllable segmentation and reduces speaker-identification accuracy on VoxCeleb1 to 26.6%, compared with 67.2% for HuBERT (Komatsu et al., 2024).

6. Applications, adaptations, and limitations

HuBERT’s masked hidden-unit principle has been extended beyond conventional speech recognition. In MelHuBERT, the learned waveform convolutional front end is replaced with 40-dimensional globally normalized log-Mel features. The 20-ms variant concatenates two adjacent 10-ms Mel frames, retains a 12-layer Transformer, uses ordinary cross-entropy, and predicts two target labels per concatenated pair. On the reported 360-hour LibriSpeech comparison, MelHuBERT-20ms uses 4.93G MACs per second versus 7.42G for HuBERT and reduces reported pretraining time by 31.2%, while achieving better phone recognition and ASR results but worse speaker-identification error in the stage-1 comparison (Lin et al., 2022).

Pac-HuBERT adapts the same masked-prediction principle to music source separation. It replaces speech MFCC-derived targets with primitive auditory features obtained from HPSS, REPET, REPET-SIM, FT2D-M, FT2D-R, and Melodia. The features are extracted from the 890-hour FMA-Large collection and clustered into 960 units. Pac-HuBERT improves source-to-distortion ratio over non-self-supervised Res-U-Net models on MusDB18, and its benefits remain present when only 25% of the supervised separation data are used (Chen et al., 2023).

For unsupervised speech segmentation and lexicon learning, HuBERT features replace CPC representations in a duration-penalized dynamic-programming system. The seventh HuBERT-Base layer is used for phone-like acoustic-unit discovery, while averaged ninth-layer HuBERT features represent hypothesized word spans. K-means clustering of these acoustic word embeddings produces an explicit lexicon. The method achieves the best reported lexicon normalized edit-distance results among the compared full-coverage systems on all five ZeroSpeech Track 2 languages, although the same English HuBERT-Base model is used cross-lingually (Kamper et al., 2024).

HuBERT has also been adapted to mixed speech. MT-HuBERT forms mixtures of clean utterances and predicts the union of their clean acoustic units using a multi-label sigmoid output and binary cross-entropy rather than a single categorical target. In 15-shot Google Speech Commands experiments, MT-HuBERT with mix-training adaptation obtains 93.80% clean Top-1 accuracy, 79.78% 2-mix Top-2 accuracy, and 65.91% unseen 3-mix Top-3 accuracy (Yuan et al., 9 Nov 2025). The approach is intended to make acoustic representations compositional across overlapping sources.

Several limitations recur across these studies. HuBERT’s pseudo-labels may encode acoustic, speaker, channel, or noise variation in addition to linguistic content. Its multiple-stage clustering process is computationally expensive. Results depend on the pretraining corpus, cluster inventory, selected Transformer layer, masking schedule, model scale, and downstream adaptation strategy. Many evaluations use English read speech, LibriSpeech, or task-specific datasets, so behavior under conversational, multilingual, noisy, spontaneous, and strongly domain-shifted conditions is not established uniformly.

HuBERT is also not synonymous with linguistic understanding. Probe performance can reflect annotation quality and probe capacity; correlation with word, speaker, gender, or semantic labels does not establish explicit symbolic knowledge. Likewise, recurring clusters in non-human vocalizations should not automatically be interpreted as phonemes or words. A canine-vocalization study uses HuBERT to discover recurring acoustic units and n-grams, but reports substantial noise categories and does not establish semantic meanings, grammar, or a validated canine language (Li et al., 2024).

The principal conceptual legacy of HuBERT is therefore methodological rather than architectural: unlabeled sequential signals can be converted into a masked prediction problem by constructing temporally aligned discrete pseudo-targets. The target-construction mechanism, the granularity of supervision, the relationship between masked pretraining and unmasked inference, and the layer from which representations are extracted all materially shape the resulting representation. HuBERT’s later extensions—multi-resolution modeling, self-distillation, speaker disentanglement, global topic supervision, multi-label mixture prediction, and domain-specific auditory clustering—retain this central principle while adapting its units and objectives to different temporal scales, signal domains, and downstream requirements.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (16)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to HuBERT.