SingMOS-Pro Benchmark for Singing Quality Assessment
- The paper introduces SingMOS-Pro, a benchmark dataset for automatic singing quality assessment with detailed MOS annotations.
- It expands on the original SingMOS preview by adding lyrics and melody ratings, diverse singing tasks, and batch-aware annotation protocols.
- Empirical results show that singing-specific models using wav2vec 2.0 and pitch features outperform speech MOS predictors in both utterance-level and system-level evaluations.
Searching arXiv for SingMOS-Pro and closely related SingMOS papers to ground the article in current literature. SingMOS-Pro is a benchmark dataset and experimental framework for automatic singing quality assessment (SQA). It is designed for training and fair comparison of models that predict mean opinion scores (MOS) of synthesized and real singing across multiple tasks and quality dimensions. Building on the earlier SingMOS preview, it expands the scale, task coverage, and annotation richness of singing MOS research by providing overall MOS for all clips and lyrics and melody MOS for a substantial subset, while also addressing the practical problem that MOS data collected in different batches can follow different subjective standards (Tang et al., 2 Oct 2025).
1. Genealogy and scope
SingMOS-Pro is explicitly positioned as an extension of the earlier SingMOS dataset. The preview version focused on SVS and SVC and provided only overall MOS per clip; it also served as the dataset for the ASRU 2024 VoiceMOS Challenge, Singing Track. SingMOS-Pro is described as a superset of that preview: it includes all SingMOS clips, adds SVR and song generation systems, introduces ground-truth recordings, and augments 4,155 clips with lyrics and melody ratings in addition to overall MOS (Tang et al., 2 Oct 2025).
This expansion is significant because the original SingMOS addressed a concrete gap in the literature: singing lacked a public MOS-annotated benchmark comparable to speech-domain resources used in VoiceMOS-style work. The preview dataset contained 3,421 singing clips, 4.25 hours of audio, Chinese and Japanese coverage, and 21,530 individual ratings, and it established a standard train/dev/test protocol for utterance-level and system-level singing MOS prediction (Tang et al., 2024). SingMOS-Pro retains that benchmarking role while broadening the problem from overall perceptual quality alone to a more fine-grained view of singing quality.
The benchmark therefore spans two related but distinct research agendas. First, it is a dataset for supervised regression of MOS from waveform-derived representations. Second, it is a testbed for studying cross-domain and cross-standard robustness, since annotation batches differ slightly in evaluation policy and subjective calibration. This dual role places SingMOS-Pro closer to a general SQA benchmark than to a single-task corpus.
2. Corpus composition
SingMOS-Pro contains 7,981 mono clips with an average length of 5.03 seconds. The total duration is approximately 11.15 hours. Sampling rates are heterogeneous: 6,937 clips are at 16 kHz, 631 clips at 24 kHz, and 413 clips at 44.1 kHz. The annotation volume is correspondingly large: 44,247 overall MOS ratings, 23,475 lyrics ratings, and 23,475 melody ratings, produced by 78 annotators (Tang et al., 2 Oct 2025).
The dataset is organized around four task categories and a system notion defined as a specific dataset-model-setting combination.
| Category | Clips | Systems / notes |
|---|---|---|
| SVS | 3,425 | 60 SVS systems |
| SVC | 1,307 | 17 SVC systems |
| SVR | 2,671 | 52 SVR systems |
| Ground truth | 578 | 12 systems |
In model terms, the benchmark covers 41 underlying models and 141 systems in total. The SVC portion includes So-VITS-SVC and NUSVC. The SVR portion combines codec families such as DAC, Encodec, and SoundStream with decoder/vocoder models including HiFi-GAN, MelGAN, Parallel WaveGAN, DiffWave, WaveGrad, and WaveNet. The SVS portion includes RNN-based sequence-to-sequence systems, XiaoiceSing, DiffSinger, VISinger, VISinger2, VISinger2+, TokSing, SingOMD, StyleSinger, TCSinger, TechSinger, SPSinger, XSinger, svsAug, and NNSVS-toolkit models such as Sinsy, ARNN-SVS, and DiffNNSVS. Closed-source demo systems and song-generation systems such as ACE-Step, DiffRhythm, Hailuo, Suno, and YuE are also represented (Tang et al., 2 Oct 2025).
The source material comprises 167 songs drawn from test sets of several public datasets and additional corpora. The benchmark uses material from CtrSVDD, adds ACE-Opencpop, ACE-KiSing, Chinese GTSinger, Japanese singing databases including Ameboshi Cipher Utagoe DB, Natsume Singing, and Namine Ritsu Utagoe DB, and also includes an auto-constructed dataset in which melodies were extracted with ChatMusician and lyrics generated with DeepSeek-V2. The resulting coverage is multilingual, at least Chinese and Japanese, and mainly in pop styles, with variation in singers, timbres, and generation paradigms (Tang et al., 2 Oct 2025).
The clip-level MOS distribution for generated samples is described as roughly Gaussian and centered between 3 and 4, whereas ground-truth clips are mostly in the 4 to 5 range, with a few noisy samples around 1. This distributional structure matters because it implies that benchmark performance cannot be interpreted as uniform across the score range; ranking near the upper end is likely especially sensitive to subtle artifacts.
3. Annotation protocol and score semantics
All listening tests are conducted online. Annotators are described as experienced listeners and received pre-annotation training; they are instructed to work in quiet environments. Every clip receives at least five ratings. For the 4,155 clips with fine-grained labels, each of the three dimensions—overall, lyrics, and melody—has at least five ratings, yielding 23,475 lyrics and 23,475 melody judgments in addition to the overall scores (Tang et al., 2 Oct 2025).
The rating scale is a 5-point Likert MOS with integer scores from 1 to 5. The three annotated dimensions have distinct semantics. Overall MOS measures the overall impression of the clip’s singing quality. Lyrics MOS targets clarity and accuracy of pronunciation and articulation. Melody MOS targets naturalness and harmony of the melodic contour, including whether the singing sounds in-tune and musical. At the clip level, the benchmark follows standard MOS aggregation:
where denotes the th score for clip .
A distinctive feature of SingMOS-Pro is that the data was collected in five batches with somewhat different purposes and instructions. Batch 1 contains only overall MOS and corresponds to the original SingMOS preview. Batches 2, 3, and 5 contain overall, lyrics, and melody MOS. Batch 4 contains only overall MOS. The paper therefore treats the benchmark as containing multiple MOS “standards”: the numeric range is shared, but the subjective interpretation of a given score can shift across batches (Tang et al., 2 Oct 2025).
Quality control is implemented through trap clips and golden clips. Trap clips contain noise or silence and are expected to receive low scores; golden clips are high-quality samples expected to receive high scores. If an annotator gives high scores to traps or low scores to goldens, the entire batch is re-annotated. This procedure is central to the benchmark’s claim of annotation reliability, especially because the benchmark aggregates ratings across multiple collection batches rather than a single homogeneous listening campaign.
4. Benchmark design and training strategies
The official split construction is batch-aware. For batches 1 to 3, systems with more than 50 clips are split 70% train and 30% test; systems with 10 to 50 clips are assigned entirely to test; systems with fewer than 10 clips are assigned entirely to train. For batches 4 and 5, all samples are used only for training. This yields three separate test sets because annotation standards differ across batches: train with 4,453 clips, test1 with 1,070 clips, test2 with 1,444 clips, and test3 with 339 clips (Tang et al., 2 Oct 2025).
For wav2vec 2.0-based SSL benchmarking, the authors restrict evaluation to 16 kHz audio because wav2vec 2.0 only supports 16 kHz. Under that restriction, the benchmark subset becomes train with 4,007 clips from 87 systems, test1 with 2,091 clips from 31 systems, test2 with 1,540 clips from 41 systems, and test3 with 376 clips from 15 systems (Tang et al., 2 Oct 2025).
The benchmark compares several model families. Speech-domain MOS predictors include DNSMOS and UTMOS. Singing-oriented baselines include SingMOS, trained on the preview dataset only, and SHEET-ssqa, trained on speech MOS datasets and the SingMOS preview. The paper’s own baselines use a wav2vec 2.0 SSL backbone with a small regression head, optionally augmented by pitch features: SSL + PM uses MIDI pitch and pitch variance, while SSL + PH uses pitch histograms. The regression loss is L1:
with SGD, learning rate 0.001, momentum 0.9, batch size 15, and 200 epochs (Tang et al., 2 Oct 2025).
Two training strategies are introduced to exploit heterogeneous MOS batches. Multi-dataset fine-tuning (MDF) first trains on the training subset of Batch 1 for 10 epochs and then fine-tunes on the entire training set. Domain ID conditioning gives the model an explicit batch or domain identifier so that prediction becomes . Conceptually, MDF acts as curriculum-based standard alignment, while Domain ID acts as explicit conditioning on annotation context. The benchmark evaluates both utterance-level and system-level RMSE, LCC, and SRCC.
5. Empirical findings
A central empirical result is that speech MOS predictors do not transfer well to singing. DNSMOS attains an utterance-level average SRCC of 0.33 and a system-level average SRCC of 0.41, while UTMOS reaches 0.36 and 0.54 respectively. Both are clearly below models trained directly on singing data, indicating that singing-specific qualities such as melody, intonation, and musical phrasing are not captured sufficiently by speech-only MOS models (Tang et al., 2 Oct 2025).
The benchmark also exposes overfitting to narrow singing domains. A model trained only on the SingMOS preview achieves utterance-level SRCC of 0.70 on test1 and system-level SRCC of 0.95 on test1, but drops to 0.09–0.48 utterance-level SRCC and 0.35–0.63 system-level SRCC on test2 and test3. This pattern matches the structural difference between the preview and the full benchmark: the earlier resource emphasized SVS and SVC with overall MOS only, whereas SingMOS-Pro introduces SVR, song generation, ground truth, broader corpus diversity, and multiple annotation standards (Tang et al., 2024).
Among the paper’s training strategies, Domain ID and MDF produce consistent improvements over naïve pooled training. Plain SSL yields an utterance-level average SRCC of 0.50 and system-level average SRCC of 0.77. Domain ID alone remains at 0.50 utterance-level average SRCC and 0.74 system-level average SRCC. MDF alone reaches 0.51 and 0.76. Domain ID plus MDF gives the best utterance-level average SRCC, 0.52, and is described as the best overall configuration for balancing utterance-level and system-level performance (Tang et al., 2 Oct 2025).
Pitch augmentation helps only modestly. SSL + PM produces an utterance-level average SRCC of 0.50 and system-level average SRCC of 0.76. SSL + PH yields approximately 0.51 utterance-level average SRCC and approximately 0.79 system-level average SRCC, the best system-level ranking among the tested methods. The paper therefore treats melody modeling as still unresolved: simple pitch features are useful but insufficient for a full account of melodic naturalness (Tang et al., 2 Oct 2025).
Follow-on work on the earlier SingMOS benchmark suggests that the modeling space extends beyond wav2vec 2.0 fine-tuning. In particular, speaker-recognition PTMs such as x-vector and ECAPA-TDNN were reported to achieve the lowest MAE and MSE among the tested PTMs on Tang et al.’s dataset, and the BATCH fusion framework, using Bhattacharyya Distance for feature alignment, established a new SOTA there (Phukan et al., 2 Jun 2025). This suggests that SingMOS-Pro is not only a dataset expansion but also a natural substrate for re-evaluating whether speaker-discriminative or music-aware encoders scale better than generic speech SSL models when the label space includes lyrics and melody.
6. Position in the literature, limitations, and extensions
Relative to prior work, SingMOS-Pro is broader in three orthogonal senses: task coverage, annotation dimensionality, and batch heterogeneity. Speech MOS benchmarks such as those behind UTMOS, DNSMOS, and related systems are speech-centric; the preview SingMOS dataset was singing-specific but limited to overall MOS and to SVS/SVC-style coverage. SingMOS-Pro adds SVR, song generation, ground-truth clips, lyrics and melody MOS, and explicit study of multi-standard training, making it the first broadly covered, multilingual, multi-dimension SQA corpus in the supplied literature (Tang et al., 2 Oct 2025).
Its limitations are also clearly defined. Language and style coverage remain incomplete: the corpus is multilingual but still concentrated on Chinese and Japanese material and mainly pop styles. Many clips come from popular open-source toolkits and challenge datasets, which may not exhaust the space of commercial or less common singing systems. Although the dataset includes lyrics and melody labels, the reported quantitative experiments focus on overall MOS prediction, leaving dedicated lyrics- and melody-prediction tasks largely open. The handling of multiple annotation standards is practical rather than fully explicit: MDF and Domain ID improve robustness, but the paper does not learn an explicit mapping between batch-specific scales in the sense of an alignment function (Tang et al., 2 Oct 2025).
These limitations delimit a concrete research agenda. The paper identifies better use of lyrics and melody scores, multi-task prediction of all three dimensions, additional SSL backbones such as HuBERT and WavLM, joint training on speech and singing MOS data, expanded language and genre coverage, and more principled handling of batch-specific MOS scales as immediate directions. A plausible implication is that efficiency-oriented architectures from adjacent speech MOS work—such as the 1,574-parameter SALF-MOS head built over frozen wav2vec features—could become relevant for low-parameter SQA, although SingMOS-Pro itself benchmarks wav2vec 2.0-based regressors rather than that architecture (Agrawal et al., 2 Jun 2025).
In that sense, SingMOS-Pro functions simultaneously as a dataset, a benchmarking protocol, and a methodological stress test. It makes the singing MOS problem large and diverse enough that simple in-domain success no longer suffices; robust performance now requires handling domain shift, annotation shift, and the separation of overall quality from lyrics and melody quality within a single benchmarking framework.