MNV-17: Mandarin NV Speech Dataset
- MNV-17 is a Mandarin speech dataset with 17 nonverbal vocalization classes, designed for joint semantic transcription and NV recognition.
- It uses LLM-generated scripts, expert validation, and precise segmentation to ensure high-fidelity NV token extraction.
- Benchmarking shows that NV-aware training enhances ASR performance, with multi-task models achieving competitive CER and NV accuracy.
MNV-17 is a performative Mandarin speech dataset for nonverbal vocalization recognition in speech, introduced to support NV-aware automatic speech recognition in settings where lexical transcription alone is insufficient. It targets nonverbal vocalizations such as sighs, laughs, and coughs, which convey emotional and intentional cues but are largely ignored by mainstream ASR systems. The corpus has a total duration of 7.55 hours, contains 2,444 utterances, and comprises 17 distinct and well-balanced classes of common NVs; the paper states that, to the best of the authors’ knowledge, this is the most extensive set of nonverbal vocalization categories assembled in a public dataset of this kind (Mai et al., 19 Sep 2025).
1. Motivation and problem setting
MNV-17 was created in response to three limitations identified in prior NV-aware ASR work: sparse NV labels, poor data sources, and limited NV coverage. The motivating premise is that existing ASR systems excel at transcribing lexical content but are almost universally insensitive to nonverbal vocalizations embedded in speech. This is consequential because NVs carry paralinguistic information relevant to expressive speech modeling, emotion-related inference, and human-like interaction (Mai et al., 19 Sep 2025).
A central design choice is the dataset’s performative character. Rather than relying on model-based post hoc detection of nonverbal events, MNV-17 uses scripted performance, so that NV instances are intentionally produced within linguistic context. The paper argues that this yields high-fidelity, clearly articulated NV tokens and substantially reduces the annotation noise associated with model-detected corpora. This also explains why class balance, rather than natural frequency, governs the corpus composition. A plausible implication is that MNV-17 is optimized as a controlled benchmark for joint lexical and nonverbal recognition rather than as a frequency-faithful sample of spontaneous conversational speech.
2. Corpus design and recording protocol
The construction process begins with script preparation. The authors surveyed prior NV-aware datasets and selected 17 distinct NV categories. Each class had equal representation, unconstrained by natural frequency. Scripts were designed to contain 1, 2, or 3 NVs per utterance in a 5:3:2 ratio, with contextually natural placement at the beginning, middle, or end of the utterance, and with NV co-occurrence explicitly allowed (Mai et al., 19 Sep 2025).
The script-writing pipeline used a LLM to generate contextually plausible and diverse sentences, after which two linguists validated all sentences for naturalness and semantic diversity. This combination of LLM-assisted generation and expert review is one of the corpus’s defining methodological features.
Recordings were made by 49 native Mandarin speakers, comprising 25 female and 24 male speakers, with mean age 26, drawn from 18 provinces in China. The participants were non-actors. Audio was recorded in a professional soundproof room, in 44.1 kHz mono, using Praat. Each speaker read approximately 50 random sentences shown on slides. No examples were provided, and speakers were encouraged to perform NVs naturally according to their own interpretation. The dataset is therefore performative but not imitation-based in the narrow sense of reproducing a fixed target rendering.
3. Annotation, segmentation, and corpus structure
A technical challenge arose from the recording format: sessions produced long audio tracks containing multiple utterances separated by short silences, making traditional Voice Activity Detection or forced alignment insufficient. To address this, the authors used an audio-capable LLM to generate precise utterance-level timestamps, and then used an ASR model to check and filter the segmented samples. Only samples with Character Error Rate below a threshold were retained. The paper presents this as a semantically and acoustically informed segmentation strategy rather than a purely signal-thresholding pipeline (Mai et al., 19 Sep 2025).
The resulting corpus organization is speaker-centric, with each speaker assigned a unique directory. The utterance distribution reflects the scripted design:
| Property | Value |
|---|---|
| Total duration | 7.55 hours |
| Total utterances | 2,444 |
| Speakers | 49 |
| One-NV utterances | 1,272 (52.0%) |
| Two-NV utterances | 715 (29.3%) |
| Three-NV utterances | 457 (18.7%) |
| Class balance | Max/min ratio 2.7 |
The max-to-min class frequency ratio of 2.7 is emphasized as a distinguishing feature. In the comparative table reported in the paper, this is substantially more balanced than the values listed for NonVerbalSpeech-38K and NonVerbalTTS. The paper also states that MNV-17 offers the broadest NV coverage among the compared corpora, with 17 classes versus 10 or fewer in the other listed datasets.
4. Nonverbal vocalization representation and labeling assumptions
MNV-17 is designed for joint semantic transcription and NV recognition rather than for isolated event classification alone. In the benchmark formulation, each NV occurrence is labeled in the transcript and counted as a single character. This makes Character Error Rate a unified metric over lexical tokens and NV tags, forcing the ASR model to solve content transcription and nonverbal recognition simultaneously (Mai et al., 19 Sep 2025).
The paper states that the dataset covers 17 balanced NV types and that these include both common and rarer nonverbal vocalizations. The provided extract does not enumerate the full inventory, but it explicitly names sighs, laughs, and coughs as examples. Because the scripts allow one, two, or three NVs per utterance, the task involves not only detection of NV presence but also recovery of type, multiplicity, and temporal order.
This design addresses a common simplification in earlier work, where datasets were limited to a few highly frequent classes such as laughter or breath. MNV-17 instead treats NV recognition as a structured sequence problem embedded in spoken language. This suggests that the dataset is intended to probe a richer hypothesis space than conventional paralinguistic tagging benchmarks.
5. Benchmarking on mainstream ASR architectures
The paper benchmarks MNV-17 on four mainstream ASR architectures: Paraformer, SenseVoice, Qwen2-Audio, and Qwen2.5-Omini. Paraformer and SenseVoice are described as non-autoregressive, with SenseVoice pretrained on multi-modal audio tasks. Qwen2-Audio and Qwen2.5-Omini are autoregressive; both were LoRA-finetuned, and Qwen2.5-Omini is described as the largest model with multi-task pretraining (Mai et al., 19 Sep 2025).
For the joint task, CER is computed over transcripts that include NV tags. The reported results are as follows:
| Model | CER (%) | NV Accuracy (%) |
|---|---|---|
| Paraformer | 5.70 | 28.64 |
| SenseVoice | 8.71 | 57.29 |
| Qwen2-Audio | 4.84 | 56.28 |
| Qwen2.5-Omini | 3.60 | 57.29 |
The paper defines NV-only recognition accuracy strictly: an utterance is counted correct only if the type, count, and order of all NVs match the ground truth exactly. Under this criterion, SenseVoice and Qwen2.5-Omini achieve 57.29%, Qwen2-Audio 56.28%, and Paraformer 28.64%.
The authors also evaluate whether NV-aware training degrades ordinary lexical ASR by comparing models trained with and without NVs in the output, then stripping NV tags from predictions and recomputing CER. The reported values are:
| Model | Non-NV model CER | NV-aware model CER |
|---|---|---|
| SenseVoice | 7.01 | 7.48 |
| Paraformer | 1.66 | 2.88 |
| Qwen2-Audio | 3.05 | 2.60 |
| Qwen2.5-Omini | 1.53 | 1.72 |
These numbers support two observations stated in the paper. First, multi-task-pretrained models substantially outperform vanilla ASR on NV recognition. Second, large multi-task and autoregressive models barely suffer degradation in lexical ASR when NV recognition is added; in the case of Qwen2-Audio, NV-aware training improves the stripped CER from 3.05 to 2.60. The paper suggests that NV symbols may provide useful context cues for surrounding words.
6. Research significance, release plan, and nomenclature
MNV-17 is positioned as a benchmark-ready resource for expressive ASR, speech synthesis, dialogue systems, and emotion-aware agents. The paper states that it includes recommended speaker-independent train, validation, and test splits, and that the dataset and pretrained model checkpoints will be made publicly available (Mai et al., 19 Sep 2025).
Its principal technical significance lies in combining four properties that are usually not achieved simultaneously in NV-aware corpora: scripted performance, human validation, semantically aware segmentation, and strong class balance across a comparatively large NV inventory. The dataset therefore functions both as a training resource and as an evaluation substrate for joint lexical-paralinguistic modeling.
A possible misconception is that MNV-17 is simply an NV classification dataset. The benchmark design contradicts that interpretation: the core task is joint semantic transcription and NV classification under sequence constraints. Another possible misconception is that it is a naturally occurring conversational corpus. The paper’s methodology instead makes clear that it is a performative dataset constructed to maximize label fidelity and balance.
The label “MNV-17” is unrelated to the use of “MNV-17 case” in recent work on Dittert’s conjecture, where it denotes the threshold in a proof valid for dimensions (Pang, 1 Jun 2026). In speech research, MNV-17 refers specifically to the Mandarin nonverbal vocalization dataset introduced for NV-aware ASR (Mai et al., 19 Sep 2025).