Existence of per-word emotion-annotated speech data

Determine whether annotated training data currently exists in which emotions are labeled separately for each word, rather than for entire sentences or passages, to support per-word emotion tagging in Deaf-centric text-to-speech systems.

Background

The paper identifies per-word emotion tagging as a participant-driven requirement for Deaf-centric text-to-speech interfaces. Implementing this capability would require training data that associates emotional labels with individual words rather than assigning a single emotional annotation to an entire sentence or passage. The availability of such data is explicitly unresolved, and its absence would constitute a major technical barrier to fine-grained, non-auditory control of emotional speech output.

References

Per-word, as opposed to per-sentence, emotion tagging requires annotated training data that has emotions annotated separately for each word, rather than entire sentences or passages. It is unclear whether such data currently exists.

Seeing the Voice, Preserving the Self: A Participatory Design Approach to Deaf-Centric Text-to-Speech  (2609.10199 - Atemnkeng et al., 9 Sep 2026) in Section 5, “Technical Requirements for Implementing the Designs,” subsection “Emotion Customization”