HArnESS: Arabic Self-Supervised Speech Models
- HArnESS is an Arabic-centric, self-supervised foundation model family comprising a bilingual teacher and compressed student variants for ASR, DID, and SER.
- It employs a HuBERT-style masked prediction framework with iterative self-distillation and PCA-based supervision compression to enhance representation quality.
- Empirical results demonstrate HArnESS-L's superior performance over HuBERT-L and XLS-R, offering a practical balance between accuracy and efficiency for diverse Arabic tasks.
HArnESS is an Arabic-centric self-supervised speech (SSL) foundation model family designed to learn speech representations that are especially useful for Arabic while remaining small enough to be practical in resource-constrained settings. Introduced as “HARNESS: Lightweight Distilled Arabic Speech Foundation Models,” it follows a HuBERT-style masked-prediction SSL regime, begins with a large bilingual Arabic-English teacher, and progressively distills that teacher into compact student models for Automatic Speech Recognition (ASR), Dialect Identification (DID), and Speech Emotion Recognition (SER) (Sukhadia et al., 31 Mar 2026). An earlier version described the acronym as “HuBERT-based Arabic and English Self-Supervised Speech,” but retained the same central program: Arabic-focused pretraining, iterative self-distillation, and compressed student variants (Sukhadia et al., 18 Sep 2025).
1. Motivation and problem setting
HArnESS was developed in response to two coupled constraints. First, Arabic speech is linguistically and acoustically diverse: Arabic spans 22 countries and many dialects; dialects differ in phonetics, morphology, and vocabulary; and real-world Arabic speech often includes English/French borrowings and code-switching (Sukhadia et al., 31 Mar 2026). Second, existing large SSL models are difficult to deploy in resource-limited settings because they require substantial compute, memory, and latency, while prior multilingual or English-centric models are not optimized for Arabic-specific variation (Sukhadia et al., 31 Mar 2026).
Within that setting, HArnESS is positioned as Arabic-centric, lightweight, and foundation-model-like. It is Arabic-centric because pretraining uses a large amount of Arabic speech, includes dialectal variety from 15 Arabic-speaking countries, uses Arabic-only distillation in the compression stage, and is benchmarked specifically on Arabic tasks: ASR, DID, and SER (Sukhadia et al., 31 Mar 2026). It is lightweight because the large teacher is distilled into structurally compressed students with much fewer parameters. It is foundation-model-like because it is pretrained on unlabeled audio and then evaluated across multiple downstream tasks rather than introduced as a task-specific acoustic model (Sukhadia et al., 31 Mar 2026).
The model family also differs from prior Arabic SSL usage in several ways. It is trained from scratch rather than adapted from a general-purpose pretrained model; it is Arabic-centric rather than primarily English-centric or broadly multilingual; it is iteratively self-distilled into compact students instead of stopping at one large model; and it uses PCA-compressed supervision to simplify teacher targets for smaller students (Sukhadia et al., 31 Mar 2026). The principal comparative baselines in the papers are HuBERT-Large and XLS-R (Sukhadia et al., 31 Mar 2026).
2. Model family and architectural organization
HArnESS is a teacher-student family built around a HuBERT-style encoder. The architecture uses a 7-layer CNN feature extractor, followed by a Transformer encoder, with a linear prediction head over cluster IDs (Sukhadia et al., 31 Mar 2026). The earlier version specifies the CNN feature encoder as having 7 temporal convolution layers with strides 5, 2, 2, 2, 2, 2, 2, kernel widths 10, 3, 3, 3, 3, 2, 2, and CNN channels 512, with all model variants using a projection dimension of 768 (Sukhadia et al., 18 Sep 2025).
| Model | Transformer specification | Parameters |
|---|---|---|
| HArnESS-L | 24 layers, embedding dim 1024 | 316M |
| HArnESS-S | 4 layers, embedding dim 1024 | 65M |
| HArnESS-ST | 4 layers, embedding dim 512 | 28M |
In the earlier formulation, the Transformer stack is further specified as follows: HArnESS-L uses FFN dimension 4096 and 16 attention heads; HArnESS-S uses FFN dimension 2048 and 16 heads; HArnESS-ST uses FFN dimension 2048 and 16 heads (Sukhadia et al., 18 Sep 2025). The parameter counts align with the paper’s compression framing: HArnESS-S corresponds to about 79.4% structural compression relative to HArnESS-L, and HArnESS-ST corresponds to about 93.7% structural compression (Sukhadia et al., 31 Mar 2026).
The family structure encodes a deliberate asymmetry. HArnESS-L is the large teacher used to learn richer bilingual representations at scale, while HArnESS-S and HArnESS-ST are compressed students intended to preserve Arabic-relevant acoustic and paralinguistic information under substantial reductions in depth, width, and attention capacity (Sukhadia et al., 31 Mar 2026). A plausible implication is that the family is meant not only as a ranking of model sizes, but as a deployment spectrum spanning high-accuracy and high-efficiency operating points.
3. Iterative self-distillation and supervision compression
HArnESS follows a multi-stage iterative self-distillation pipeline. In the first stage, MFCC features are extracted from raw speech, clustered with K-means to create pseudo-labels, and used to train the first model. In the second stage, layer 9 representations from the first model are clustered to generate improved pseudo-labels for a stronger large model. In the third stage and beyond, last-layer embeddings from the previous model are reclustered, and compression is introduced through student architectures (Sukhadia et al., 31 Mar 2026). The earlier version states that the system uses three iterations, with pseudo-label generation based on 39-dimensional MFCC features in iteration 1, layer 9 embeddings in iteration 2, and last-layer embeddings of the second-iteration HArnESS-L model in iteration 3 (Sukhadia et al., 18 Sep 2025).
The compression schedule is explicit. In the early iterations the architecture stays unchanged to learn better abstractions, and starting from the third iteration the model is compressed along three axes: depth reduction, width reduction, and attention reduction (Sukhadia et al., 31 Mar 2026). The teacher is pretrained on a 23K-hour Arabic-English corpus in the first two iterations, while the lightweight students are distilled on about 1100 hours of Arabic audio from QASR training data in the third iteration (Sukhadia et al., 18 Sep 2025). The earlier version further reports 500k training steps for iteration 1, 700k for iteration 2, and 300k for iteration 3, with K-means trained on a 300-hour subset at the teacher stage and on a random 30% subset of the Arabic-only student-stage data (Sukhadia et al., 18 Sep 2025).
A notable contribution is PCA-based compression of teacher supervision. Instead of clustering the full teacher embedding , the paper optionally projects it to a smaller space with , and then performs clustering on (Sukhadia et al., 31 Mar 2026). The papers are explicit that PCA compresses the teacher supervision signal, not the student input, and that it is used to remove redundant or noisy directions, produce cleaner clustering targets, reduce supervision complexity, and better match the capacity of shallow and thin students (Sukhadia et al., 31 Mar 2026). They also report that PCA-reduced supervision leads to faster convergence and more stable optimization (Sukhadia et al., 31 Mar 2026).
The training objective is HuBERT-style masked prediction. Some frames are masked, the model predicts discrete pseudo-labels for those frames, and training is done with cross-entropy over cluster IDs (Sukhadia et al., 31 Mar 2026). The loss is computed on both masked and unmasked frames with fixed weighting; the earlier version describes the effective form as
where both terms are cross-entropy losses over discrete pseudo-labels (Sukhadia et al., 18 Sep 2025). The stated rationale is that masked frames encourage contextual reasoning, while unmasked frames stabilize training and prevent collapse (Sukhadia et al., 31 Mar 2026).
4. Evaluation protocol and downstream tasks
HArnESS is evaluated on three representative Arabic speech tasks. ASR is fine-tuned on a 300-hour QASR subset and tested on MGB2 and MGB3, with word error rate (WER) as the metric. DID is evaluated on ADI5 with five dialect classes—MSA, Egyptian, Levantine, North African, and Gulf—and uses accuracy as the metric. SER is evaluated on KSUEmotion with six emotion classes and also uses accuracy (Sukhadia et al., 31 Mar 2026).
The evaluation design distinguishes content recognition from representation quality. For DID and SER, the SSL encoder is frozen and only a small classifier is trained on top, so the results reflect the quality of the learned representations rather than end-to-end task-specific adaptation (Sukhadia et al., 31 Mar 2026). The earlier version adds that DID and SER use average embeddings from all SSL layers with a small CNN + self-attention classifier, and that ASR fine-tuning uses an encoder-decoder model with joint CTC + attention loss in ESPnet, with a 2-conformer-layer encoder and a 2-transformer-layer decoder (Sukhadia et al., 18 Sep 2025).
The benchmark framing is comparative rather than isolated. HArnESS is explicitly tested against HuBERT-L and XLS-R, and the 2025 version also reports task-specific SOTA upper bounds for contextualization: ASR MGB2 at 10.24 WER, ASR MGB3 at 21.31 WER, SER at 83.31%, and DID at 82.5% (Sukhadia et al., 18 Sep 2025). This situates the HArnESS family simultaneously against general-purpose SSL baselines and against stronger task-specific systems that may use more specialized supervision.
5. Reported results and compression trade-offs
The principal empirical finding is that HArnESS-L outperforms HuBERT-L and XLS-R consistently on the evaluated Arabic tasks (Sukhadia et al., 31 Mar 2026). On ASR for MGB2, the reported WERs are 22.6 for HuBERT-L, 22.60 for XLS-R, and 15.50 for HArnESS-L; on MGB3, 51.2 for HuBERT-L, 51.80 for XLS-R, and 41.60 for HArnESS-L. On SER, the reported accuracies are 91.92% for HuBERT-L, 73.32% for XLS-R, and 94.66% for HArnESS-L. On DID, the corresponding accuracies are 64.14%, 42.35%, and 84.98% (Sukhadia et al., 31 Mar 2026).
These comparisons support two claims made in the papers. Versus HuBERT-Large, HArnESS-L outperforms an English-trained SSL model on Arabic downstream tasks. Versus XLS-R, HArnESS-L also outperforms a large multilingual SSL model consistently (Sukhadia et al., 31 Mar 2026). This suggests that Arabic-centric pretraining is more effective than using an English-centered SSL model for Arabic speech, and that Arabic-centric pretraining can outperform generic multilingual scale on the reported tasks (Sukhadia et al., 31 Mar 2026).
Compression introduces clear but nonuniform degradation. HArnESS-S, at about 79.4% structural compression, reports 20.20 WER on MGB2, 52.80 WER on MGB3, 91.15% on SER, and 70.84% on DID. HArnESS-ST, at about 93.7% structural compression, reports 23.20 WER on MGB2, 58.20 WER on MGB3, 89.02% on SER, and 69.77% on DID (Sukhadia et al., 31 Mar 2026). The papers repeatedly note that the compressed students often remain competitive despite substantial reduction, but that DID is especially sensitive to compression (Sukhadia et al., 31 Mar 2026).
The 2025 version makes the trade-offs more explicit. An additional width-reduction variant yields even smaller models but performs substantially worse, especially on dialect identification; reducing embedding dimension too aggressively to produces 96.52% compression but sharply worsens accuracy and WER (Sukhadia et al., 18 Sep 2025). Attention-head reduction from 16 to 4 yields an additional 26.15% compression and hurts DID more than ASR or SER (Sukhadia et al., 18 Sep 2025). The reported ablations therefore indicate that embedding reduction is the most damaging compression knob, attention reduction is less harmful but still costly for DID, and initialization choice has little effect in the final student stage (Sukhadia et al., 18 Sep 2025).
A further point from the earlier version is that HArnESS-L surpasses the task-specific SOTA reference on SER and DID while remaining substantially above HuBERT and XLS-R on all tasks (Sukhadia et al., 18 Sep 2025). This does not imply universal dominance across Arabic speech workloads, but it does place the large teacher in a strong regime for paralinguistic and dialectal evaluation as defined in the paper.
6. Practical significance and terminological scope
HArnESS is presented as a practical solution for resource-constrained Arabic speech applications because it gives users an explicit operating choice: use HArnESS-L for best accuracy, HArnESS-S for a better memory/latency balance, or HArnESS-ST for a very compact model (Sukhadia et al., 31 Mar 2026). The papers identify edge devices, low-memory servers, real-time Arabic speech tools, dialectal ASR systems, emotion-aware conversational systems, and on-device language technologies in Arabic-speaking regions as especially relevant deployment settings (Sukhadia et al., 31 Mar 2026). Because the students are much smaller than large SSL baselines, the model family is positioned as deployable where large encoders would be too expensive (Sukhadia et al., 31 Mar 2026).
Conceptually, the work’s broader contribution is to show that Arabic-centric training plus careful distillation can beat general SSL baselines like HuBERT and XLS-R on Arabic tasks while producing lighter models suitable for deployment (Sukhadia et al., 31 Mar 2026). This is not simply a matter of pruning a multilingual backbone; the family is trained from scratch, iteratively self-distilled, and shaped around Arabic-relevant acoustic and paralinguistic structure (Sukhadia et al., 31 Mar 2026).
The term “harness” has a different meaning in several contemporaneous arXiv papers, where it denotes the runtime scaffold around LLM agents, robot middleware, or execution-layer governance rather than a speech model family. Examples include “HarnessBridge,” which learns a bidirectional controller over the agent–environment interface (Wang et al., 11 Jun 2026), and “Harness Engineering for Physical AI,” which argues that robot middleware is the harness layer for control, computing, and communication (Lee et al., 8 Jun 2026). In the speech literature, however, HArnESS denotes the Arabic-centric SSL foundation-model family described above, not an orchestration layer.