EmoTale: Bilingual Speech-Emotion Corpus
- EmoTale is a bilingual enacted speech-emotion corpus containing 800 recordings across Danish and English with five emotion classes.
- The dataset uses controlled elicitation of neutral sentences from Emo-DB, providing both categorical and dimensional emotion annotations validated via leave-one-speaker-out cross-validation.
- Baseline experiments indicate that self-supervised speech model embeddings outperform hand-crafted features, establishing EmoTale as a robust SER evaluation framework.
Searching arXiv for the specified EmoTale paper and closely related work to ground the article. arxiv_search.query({"2search_query2 OR ti:\2"EmoTale: An Enacted Speech-emotion Dataset in Danish\"","start":2search_query2,"max_results":5}) arxiv_search.query({"2search_query2 Enhanced Talking Face Generation\" OR id:(&&&2id:(Hjuler et al., 20 Aug 2025) OR ti:\2&&&)","start":2search_query2,"max_results":5}) arxiv_search.query({"2search_query2 Speech-Driven Animation with Content-Emotion Disentanglement\" OR id:(Daněček et al., 2023)","start":2search_query2,"max_results":5}) EmoTale is a bilingual enacted speech-emotion corpus for Danish and English designed to provide a functional benchmark for speech emotion recognition (SER) in Danish, a language with very limited emotional speech resources. It was introduced as a response to the scarcity of Danish SER datasets and, in particular, to the limitations of Danish Emotional Speech (DES; 2id:(Hjuler et al., 20 Aug 2025) OR ti:\2997), which contains four speakers, includes single words and questions, and was not specifically curated for SER. EmoTale comprises 82search_query2search_query2^ recordings, provides both categorical and dimensional emotion annotations, and is validated through baseline SER experiments using self-supervised speech model embeddings and hand-crafted acoustic features under leave-one-speaker-out evaluation (&&&2search_query2&&&).
2id:(Hjuler et al., 20 Aug 2025) OR ti:\2. Historical position and design rationale
EmoTale was created against a narrow Danish-resource background. According to its description, DES, published in 2id:(Hjuler et al., 20 Aug 2025) OR ti:\2997, is the only other database of Danish emotional speech. DES comprises four speakers, enacts five emotions—neutral, surprise, happiness, sadness, and anger—and reported average human classification accuracy of 67.3%. EmoTale was designed to provide a more modern SER testbed with more speakers, bilingual recordings, richer annotations, and standardized evaluation protocols (&&&2search_query2&&&).
Its design choices are explicitly benchmark-oriented. Unlike DES, EmoTale uses emotionally neutral sentences translated from Emo-DB and focuses on minimizing semantic and contextual bias while supporting transferability studies. The corpus therefore emphasizes controlled elicitation and comparability rather than spontaneity. This suggests a deliberate trade-off: ecological validity is reduced relative to naturalistic emotional speech, but experimental control is improved for speaker-independent SER and cross-corpus transfer analysis (&&&2search_query2&&&).
The dataset’s novelty is defined by four elements: bilingual composition, enacted categorical labels plus dimensional ratings, validation with baseline SER models under leave-one-speaker-out cross-validation, and public release of data and code via a GitHub repository. In this respect, EmoTale functions not only as a corpus but also as a reproducible evaluation framework for Danish affective computing (&&&2search_query2&&&).
2. Corpus composition and recording protocol
EmoTale contains 452search_query2^ Danish utterances and 352search_query2^ English utterances, for a total of 82search_query2search_query2^ recordings balanced across five emotions: Neutral (N), Anger (A), Sadness (S), Happiness (H), and Boredom (B). The recordings were produced by 2id:(Hjuler et al., 20 Aug 2025) OR ti:\28 participants, of whom 2id:(Hjuler et al., 20 Aug 2025) OR ti:\22^ were female and six male. The age distribution spans 9 to 39 years, with mean age 22.8, and includes children; the study was IRB-approved and used GDPR-compliant consent (&&&2search_query2&&&).
| Component | Value | Notes |
|---|---|---|
| Languages | Danish: 452search_query2; English: 352search_query2^ | 82search_query2search_query2^ recordings total |
| Speakers | 2id:(Hjuler et al., 20 Aug 2025) OR ti:\28 | 2id:(Hjuler et al., 20 Aug 2025) OR ti:\22^ female, 6 male |
| Emotions | 5 classes | N, A, S, H, B |
Participants were recruited with priority given to acting experience and Danish/English fluency; exclusion criteria were age under 7 or lack of Danish proficiency. Recordings were made across multiple sessions in quiet rooms with no ambient noise. Participants could repeat sentences, but only the last take was retained. Audio was recorded using a RØDE Wireless Go wireless lavalier microphone at 48 kHz in WAV format, and some files were cropped to remove keyboard clicks at the start or end. Bit depth and total duration are not reported (&&&2search_query2&&&).
The verbal material consists of five emotionally neutral sentences selected and translated from Emo-DB so that Danish and English share the same semantic content. Examples include “Dugen ligger på køleskabet.” / “The tablecloth is lying on the fridge.” and “Om syv timer er det morgen.” / “In seven hours it will be morning.” Each sentence was enacted in all five target emotions. Elicitation was performed through self-induced recall of strongly felt situations, but the resulting speech remains enacted rather than spontaneous (&&&2search_query2&&&).
This controlled sentence design has methodological significance. Because all speakers render the same small set of neutral prompts across emotions and languages, lexical variability is minimized. A plausible implication is that measured performance is more attributable to prosodic and acoustic affect cues than to lexical or semantic confounds (&&&2search_query2&&&).
3. Annotation scheme and reliability
EmoTale provides both categorical and dimensional annotation. The categorical taxonomy consists of the five classes Neutral, Anger, Sadness, Happiness, and Boredom. The dimensional framework assigns arousal, valence, and dominance ratings on a 2id:(Hjuler et al., 20 Aug 2025) OR ti:\2–5 scale with 2search_query2.5 increments. The definitions are explicit: valence ranges from 2id:(Hjuler et al., 20 Aug 2025) OR ti:\2^ = negative to 5 = positive, arousal from 2id:(Hjuler et al., 20 Aug 2025) OR ti:\2^ = calm to 5 = excited, and dominance from 2id:(Hjuler et al., 20 Aug 2025) OR ti:\2^ = submissive/weak to 5 = dominant/strong (&&&2search_query2&&&).
Each utterance was labeled by three independent annotators—the first, second, and last authors—with one categorical emotion and three dimensional scores. Inter-annotator reliability for categorical labels was assessed with Cohen’s kappa, and agreement with the predefined intended emotion was also measured. Reported values are PRESERVED_PLACEHOLDER_2search_query2, PRESERVED_PLACEHOLDER_2id:(Hjuler et al., 20 Aug 2025) OR ti:\2, and , using
where is observed agreement and is chance agreement (&&&2search_query2&&&).
For dimensional labels, reliability was measured using the Concordance Correlation Coefficient (CCC). The reported CCC values are 2search_query2.72 for arousal, 2search_query2.75 for valence, and 2search_query2.57 for dominance. The dataset description characterizes arousal and valence agreement as moderate to strong, and dominance agreement as moderate. Krippendorff’s alpha was not reported (&&&2search_query2&&&).
The annotation design also supports comparison with model-derived dimensional predictions. For modeling, dimensional outputs from pretrained models were rescaled from to via
with analogous transformations for valence and dominance. This creates direct comparability between machine predictions and human annotation scales (&&&2search_query2&&&).
4. Baseline SER methodology
EmoTale was validated through a baseline SER pipeline comparing self-supervised speech model embeddings with openSMILE hand-crafted acoustic features. The self-supervised component included wav2vec 2.2search_query2^ base (facebook/wav2vec2-base), embeddings extracted from the last transformer layer and average-pooled over time, as well as two fine-tuned models: ehcalabres/wav2vec2-lg-xlsr-en-speech-emotion-recognition for categorical labels (w2v2-FT-cat) and audeering/wav2vec2-large-robust-^^^^2id:([2508.14548](/papers/2508.14548)) OR ti:\22^^^^-ft-emotion-msp-dim for dimensional labels (w2v2-FT-dim). These embeddings were fed into a linear support vector classifier, with no classifier fine-tuning inside the pretrained model (&&&2search_query2&&&).
The hand-crafted baseline used openSMILE with the eGeMAPS and ComParE feature sets, again classified with a linear SVC. The use of a linear kernel is stated to align with SER practice in ComParE challenges. Evaluation employed leave-one-speaker-out cross-validation, where each fold leaves one speaker out for testing and trains on the remaining speakers. Within each fold, features were standardized using the training-fold mean and standard deviation:
For pretrained-model compatibility, recordings were downsampled to 2id:(Hjuler et al., 20 Aug 2025) OR ti:\26 kHz and stereo channels were averaged into mono (&&&2search_query2&&&).
The primary metric was Unweighted Average Recall (UAR),
PRESERVED_PLACEHOLDER_2id:(Hjuler et al., 20 Aug 2025) OR ti:\2search_query2^
with Macro-F2id:(Hjuler et al., 20 Aug 2025) OR ti:\2^ also defined in the study. Results were reported in three forms: aggregated UAR from a single confusion matrix over all leave-one-speaker-out predictions, “Speaker UAR” as the mean of per-fold UAR values, and “Sentence UAR” as the mean UAR over sentence groups. Because the SVC has fixed parameters and no randomness, aggregated UAR has standard deviation 2search_query2^ under the reported protocol (&&&2search_query2&&&).
The methodological contribution lies not only in the use of multiple feature regimes but also in the comparison against reference corpora—Emo-DB, DES, Urdu, and AESDD—under a common pipeline. This makes EmoTale suitable for both within-corpus benchmarking and cross-corpus transfer analysis (&&&2search_query2&&&).
5. Quantitative performance and cross-corpus behavior
On EmoTale, the best reported performance was obtained by w2v2-FT-dim, which achieved 64.2id:(Hjuler et al., 20 Aug 2025) OR ti:\2% aggregated UAR, with Speaker UAR PRESERVED_PLACEHOLDER_2id:(Hjuler et al., 20 Aug 2025) OR ti:\2id:(Hjuler et al., 20 Aug 2025) OR ti:\2^ and Sentence UAR PRESERVED_PLACEHOLDER_2id:(Hjuler et al., 20 Aug 2025) OR ti:\22. The other reported results on EmoTale were 59.6% for w2v2-FT-cat, 52.2search_query2% for ComParE, 46.2search_query2% for eGeMAPS, and 29.7% for wav2vec 2.2search_query2^ base. The paper concludes from these experiments that embeddings are superior to hand-crafted features on the corpus (&&&2search_query2&&&).
Using the same pipeline on DES, w2v2-FT-dim reached 67.7% aggregated UAR, w2v2-FT-cat 62.7%, ComParE 48.5%, eGeMAPS 42.7%, and wav2vec base 32.7%. Although DES slightly exceeds EmoTale in top-line UAR under the best model, the study notes that DES is more sentence-dependent. The explanation given is that DES contains single words and questions such as “Nej” / “No,” where question intonation may make emotion recognition easier, whereas EmoTale uses neutral sentence enactments (&&&2search_query2&&&).
Cross-lingual experiments were conducted on four shared emotions—happy, angry, sad, and neutral. The reported conclusion is that models trained on EmoTale generalize comparably to those trained on other corpora and that deep features transfer better across corpora than hand-crafted features. A particularly notable observation is that, when inferring on Emo-DB, cross-domain UAR can exceed in-corpus UAR for models trained on EmoTale and DES, which the authors interpret as consistent with Emo-DB’s strong perceptual filtering during dataset creation, specifically a greater-than-82search_query2% human recognition criterion (&&&2search_query2&&&).
Additional validation was carried out for dimensional labels. Predictions from w2v2-FT-dim, rescaled to the human label range, showed good agreement with human annotations for arousal and dominance, whereas valence agreement was systematically lower across datasets. The paper identifies this as an established trend in SER. This suggests that EmoTale is useful not only for discrete emotion classification but also for dimensional affect modeling, albeit with the usual valence-related difficulties (&&&2search_query2&&&).
6. Interpretation, limitations, access, and nomenclature
The study offers several interpretations of its empirical results. Fine-tuned self-supervised speech model embeddings are said to capture rich prosodic and spectral patterns linked to affect and to benefit from both large-scale pretraining and task-specific fine-tuning. By contrast, fixed low-level acoustic descriptors are less robust across speakers and corpora. EmoTale’s broader age range, inclusion of children, and bilingual design introduce speaker and language variability, yet the embeddings still show strong performance and transferability. The corpus is balanced across emotions, which reduces class-imbalance concerns, but speaker-level variability remains substantial, as indicated by the Speaker UAR standard deviation of approximately 2id:(Hjuler et al., 20 Aug 2025) OR ti:\22.4% for w2v2-FT-dim (&&&2search_query2&&&).
Several limitations are stated explicitly. The emotions are enacted rather than spontaneous; the dataset is modest in size and therefore not suitable for training large end-to-end models; and annotation biases may persist, especially for valence, while dominance agreement is only moderate. These caveats locate EmoTale primarily as an evaluation and benchmarking resource rather than a large-scale pretraining corpus (&&&2search_query2&&&).
The intended applications include use as a Danish SER benchmark, evaluation of SER for Danish speakers including children, cross-lingual and multilingual transfer studies between Danish and English and other languages, augmentation and few-shot adaptation studies, dimensional SER evaluation, and ASR-related work enabled by the controlled bilingual text content. Filenames encode language, speaker ID, emotion, and sentence index, for example DK_^^^^2search_query2search_query2^^^^4_A_5.wav, and the repository provides data and code at https://github.com/snehadas/EmoTale. Distribution is under copyright, raw data may be provided upon request, citation is requested, and maintenance is assigned to the corresponding author (&&&2search_query2&&&).
A potential source of confusion is the name itself. “EmoTale” also appears in other affective computing contexts, including an emotion-conditioned talking-face generation framework described in “Emotionally Enhanced Talking Face Generation” (&&&2id:(Hjuler et al., 20 Aug 2025) OR ti:\2&&&). Related neighboring work on emotion-controllable speech-driven facial animation includes EMOTE, a 3D method with content-emotion disentanglement (Daněček et al., 2023), and EmoTalker, a diffusion-based emotionally editable talking-face system (Zhang et al., 2024). In the present usage, however, EmoTale denotes the bilingual Danish-English enacted emotional speech corpus introduced in 22search_query225 (&&&2search_query2&&&).