Papers
Topics
Authors
Recent
Search
2000 character limit reached

Comparison Category Rating (CCR)

Updated 12 July 2026
  • Comparison Category Rating (CCR) is a subjective quality-assessment method that uses paired relative judgments to indicate improvement, equality, or degradation.
  • It employs randomized stimulus presentation and arithmetic averaging to create a robust Comparison Mean Opinion Score (CMOS) that minimizes individual scale biases.
  • Widely applied in speech enhancement and video quality assessment, CCR captures subtle perceptual differences and accommodates heterogeneous listener preferences.

Comparison Category Rating (CCR) is a subjective quality-assessment method based on paired relative judgment. A listener or viewer is presented with two stimuli that share the same content and rates the quality of the second relative to the first, so the resulting score expresses perceived improvement, equality, or degradation rather than an absolute quality level. Across speech, audio, and video assessment, CCR is used when direct comparison is preferable to absolute rating, when small perceptual differences matter, and when a system may either improve or degrade a signal relative to a reference (Naderi et al., 2021, Naderi et al., 2022).

1. Definition and relation to adjacent subjective methods

CCR belongs to the family of double-stimulus methods. In the speech literature, listeners hear two versions of the same utterance and rate the second relative to the first; in recent speech crowdsourcing and a comparative video study, the scale is a seven-category symmetric scale mapped to integers from 3-3 to +3+3, with anchors from “Much worse” to “Much better.” The P.910-oriented video crowdsourcing summary also describes a five-category CCR variant mapped from 2-2 to +2+2, again centered on “Same.” This suggests that CCR is best understood as a comparative design principle rather than a single invariant scale specification (Suárez et al., 2023, Naderi et al., 24 Sep 2025, Naderi et al., 2022, Naderi et al., 2020).

CCR is distinct from Absolute Category Rating (ACR), Degradation Category Rating (DCR), and ACR with hidden reference (ACR-HR). ACR is single-stimulus and produces an absolute Mean Opinion Score (MOS), typically on a five-point scale. DCR is also double-stimulus, but the first item is an explicit reference and the task is to rate the degradation of the second item relative to that reference. ACR-HR remains single-stimulus but inserts hidden references and converts per-viewer ACR scores into a differential score, aggregated as DMOS. By contrast, CCR directly asks whether the second stimulus is better, the same, or worse than the first, and therefore can represent both impairment and enhancement.

Method Stimulus framing Aggregate
CCR Paired relative judgment of second vs first CMOS
ACR Single-stimulus absolute judgment MOS
DCR Explicit reference followed by processed signal DMOS
ACR-HR Single-stimulus with hidden references DV / DMOS

The principal methodological rationale for CCR is that direct comparison reduces variability associated with individual use of absolute scales and increases sensitivity to subtle perceptual differences. In speech enhancement, this is particularly relevant because noise suppression often improves background-noise attenuation while simultaneously affecting speech naturalness or introducing artifacts. In video quality assessment, the same relative framing makes CCR sensitive to improvements beyond the reference, including restoration and super-resolution cases.

2. Stimulus presentation, scales, and score construction

In standard CCR practice, the two stimuli are presented in randomized order, commonly as AABB or BBAA, and the participant rates the second stimulus relative to the first. When one stimulus is the processed item and the other is the unprocessed or reference item, the stored score is sign-corrected after collection so that the final value consistently represents “processed relative to reference.” This order randomization differentiates CCR from DCR and is intended to reduce anchoring and order biases while preserving the sensitivity of paired comparison (Suárez et al., 2023, Naderi et al., 2021, Naderi et al., 2020).

In speech crowdsourcing, instructions have been used to emphasize that “overall quality” includes both distortion or noise added or removed and perceived speech quality. That instruction is consequential because CCR does not force raters to privilege either background-noise suppression or speech fidelity; rather, it allows the relative judgment to encode a trade-off between them. In one crowdsourced speech study, two user-interface designs were tested: a two-player design in which participants could replay clips in arbitrary order, and a single-player sequential design with a one-second pause between clips. The latter reduced order confusion by making the “second compared to first” framing explicit.

The core aggregation is arithmetic averaging of the signed category scores. For a stimulus pair with NN valid ratings sis_i, the Comparison Mean Opinion Score is

+3+30

In the speech-enhancement study that analyzed clip- and condition-level aggregation, each clip mean was first computed as

+3+31

and the mean per condition over +3+32 clips was then

+3+33

Uncertainty was summarized with standard variance, standard error, and +3+34-based confidence intervals:

+3+35

+3+36

These constructions make CCR directly usable as a condition-level response variable in downstream regression, correlation, and mixed-effects analysis.

3. Crowdsourcing implementations and quality-control architecture

CCR has been adapted from laboratory standards into scalable crowdsourcing frameworks for both speech and video. For speech, an open-source P.808 implementation on Amazon Mechanical Turk extended the standard ACR workflow to include DCR and CCR and integrated qualification into the main rating task. For video, an open-source P.910-Crowd framework implemented ACR, ACR-HR, DCR, and CCR and added rater, environment, hardware, and network qualifications, together with gold and trapping questions (Naderi et al., 2020, Naderi et al., 2022).

The speech-oriented P.808 adaptation organizes CCR collection into instruction, qualification, setup testing, training, and rating. Qualification can include hearing screening and demographic questions. Setup testing includes listening-level adjustment, a headset or earpod check, and an environment suitability test based on a modified Just-Noticeable Difference in Quality procedure. Training familiarizes participants with the CCR scale and can include a gold item in which both stimuli are references, reinforcing that “About the Same” corresponds to +3+37. Rating sessions commonly contain 10–12 paired stimuli per session, with at least one gold or trap item (Naderi et al., 2021).

Quality control is layered. In the P.808 implementation, submissions can be rejected if full playback is not completed, the two-eared check is failed, or the trapping item is answered incorrectly; ratings can then be excluded from analysis if the environment test fails, gold items are missed, or score variance is too low to rule out straightlining. The same implementation reports that integrated qualification reduced end-to-end execution time by 4–5+3+38 compared to a two-stage qualification-plus-rating design, and that a temporal environment certificate reduced overall worker time by about 40% (Naderi et al., 2020).

The speech-enhancement CCR study provides a concrete crowdsourcing example. Using the ITU-T P.808 toolkit on Mechanical Turk, 216 workers completed 810 assignments; after screening, 626 assignments were retained, producing 5,008 usable votes from assignments containing eight clip pairs plus one gold question. Headphone checks, environment suitability checks, a gold standard item, and participant-level variance checks were all part of the filtering pipeline. Such procedures are central because crowdsourcing introduces uncontrolled device and environment variability that laboratory CCR does not face (Suárez et al., 2023).

4. Statistical analysis, reproducibility, and preference structure

CCR data are often analyzed at both the clip and condition levels. Correlation with objective metrics is commonly summarized with Pearson and Spearman coefficients, and predictive correspondence may also be summarized by RMSE:

+3+39

2-20

In speech enhancement, condition-level aggregation produced stronger statistical correspondence to objective metrics than clip-level analysis, highlighting the stabilizing effect of averaging over multiple clips per condition. The same study modeled objective scores as functions of CMOS and condition parameters such as SNR, speech level, background type, transient amount, and a positive-versus-negative preference factor, using both fixed-effects and mixed-effects models with Bonferroni-corrected thresholds (Suárez et al., 2023).

A central empirical finding is that CCR distributions can be bimodal rather than unimodal. In the speech-enhancement study, the histogram of mean score per worker showed two clusters: one group tending to reward noise suppression even when artifacts are introduced, and another tending to penalize speech distortions or residual noise. Adding a positive/negative affiliation factor significantly improved model fit in ANOVA, with reported 2-21-values as low as 2-22. This indicates that CCR may encode heterogeneous listening strategies rather than a single latent preference axis.

Reproducibility has been strong in crowdsourced speech CCR. In one study with three independent Mechanical Turk runs on ITU-T Supplement P.23 material, per-condition average valid votes were 60.1, 69.3, and 66.2, average 95% confidence-interval widths were in 2-23, run-to-run Pearson correlations ranged from 2-24 to 2-25, and 2-26 reached 2-27. A linear mixed-effects model found a significant main effect only for degradation condition, with no run effect or interaction, indicating stability across runs (Naderi et al., 2021).

Comparable stability appears in video crowdsourcing. In the six-study video comparison, CCR test–retest reliability for synthetic impairments at condition level reached Pearson 2-28, Spearman 2-29, and Kendall +2+20 +2+21 for good-quality sources; at clip level the corresponding Pearson correlation was +2+22. These results support the use of CCR as a reproducible condition-level measure even when per-item variability is non-negligible (Naderi et al., 24 Sep 2025).

5. Domain-specific uses in speech enhancement and video quality assessment

In speech enhancement, CCR is especially suited to evaluating algorithms whose output may improve background-noise conditions while changing speech quality. A recent crowdsourced study compared processed and unprocessed clips from three real-time communications systems—Microsoft Teams, Zoom, and Amazon Chime—using noisy speech generated from VCTK utterances combined with Scaper-synthesized office, cafeteria, and park scenes over SNR values of +2+23 dB and +2+24 dB, speech levels of +2+25 dB and +2+26 dB SPL, and transient-noise occupancies of approximately 30% and 90% of clip duration. The design yielded 24 noisy conditions, 240 noisy clips, 720 enhanced signals, and 72 processed conditions. Participants reliably used the CCR scale to indicate improvements or degradations introduced by the RTC noise suppressors (Suárez et al., 2023).

That study also clarified how CCR relates to objective speech-quality metrics. Spearman correlations between condition-level CCR means and objective scores differed by preference mode. For DNSMOS-Speech, the reported correlations were 0.69 overall, 0.68 for negative-mode raters, and 0.54 for positive-mode raters. For 3QUEST-Speech, the values were 0.50 overall, 0.62 for negative-mode raters, and 0.31 for positive-mode raters. The stated implication was that positive raters align better with noise-focused metrics, whereas negative raters align better with speech-quality-focused metrics. No single objective score was found to universally substitute for CCR across modes.

In broader speech QoE benchmarking, crowdsourced CCR has been compared against ACR in both laboratory and crowdsourcing settings. On ITU-T Rec. P.863 dataset 401, the global correlation between crowdsourced CCR CMOS and laboratory ACR MOS was Pearson +2+27 and Spearman +2+28 across 48 conditions; the correlation with crowdsourced ACR MOS was Pearson +2+29 and Spearman AA0. After removing seven strong rank-order outliers, the CMOS–MOS Pearson correlation increased to AA1. The paper also reported impairment-specific divergences: discontinuity was less penalized in CCR than ACR, whereas coloration tended to be more salient in CCR (Naderi et al., 2021).

In video quality assessment, CCR has been used for synthetic impairments and bitrate-ladder design. In a side-by-side six-study comparison with ACR-HR, CCR was more sensitive to quality changes and captured improvements beyond the reference, notably for DNN super-resolution upscaling. For bitrate-ladder decisions, agreement with direct pair-comparison tests was reported as 100% for good-quality sources and 60% for fair-quality sources, compared with 40% and 20% for ACR-HR. At the same time, ACR-HR was approximately twice as fast and cost-effective and had lower normalized variability. Method choice therefore materially changed inferred saturation points and bitrate recommendations (Naderi et al., 24 Sep 2025).

Domain Representative empirical finding Source
Speech enhancement Two preference groups emerged; SNR was a dominant factor (Suárez et al., 2023)
Speech QoE crowdsourcing Run-to-run reproducibility reached AA2 (Naderi et al., 2021)
Video quality assessment CCR captured improvements beyond the reference and matched pair-comparison decisions more often than ACR-HR (Naderi et al., 24 Sep 2025)

6. Limitations, controversies, and methodological guidance

CCR is not free of bias or design sensitivity. In degradations-only scenarios, the effective CMOS range may collapse to AA3, whereas ACR MOS spans AA4; this contributes to the “banana-shaped” scatter reported when CMOS is compared with MOS. Listening to the same content twice can also change what becomes salient: discontinuities may be less disruptive because the content is already known, whereas coloration can become more obvious under direct comparison. A plausible implication is that CCR and ACR should not be treated as interchangeable even when they are strongly correlated overall (Naderi et al., 2021).

Crowdsourcing adds further complications. Residual device and environment variability remain even with strong qualification and screening. In speech enhancement, the bimodality of worker preferences increases variance and requires mode-aware analysis. In video assessment, CCR showed slightly higher normalized variability and higher time and monetary cost than ACR-HR. These observations do not negate CCR’s usefulness; rather, they indicate that sensitivity is obtained at the cost of a more demanding experimental design (Suárez et al., 2023, Naderi et al., 24 Sep 2025).

Several design recommendations recur across the literature. Randomized presentation order with post-hoc sign correction is essential. Gold or trap items using identical reference pairs should be included, with the expected response near AA5. In speech, instructions should emphasize overall quality rather than only noise reduction or only speech naturalness, and balanced speakers per condition are advisable because speaker or gender can affect objective scores even when CCR itself shows no significant speaker effect in a given study. Loudness normalization is important when algorithms apply automatic gain control, and more than 10 clips per condition is recommended where feasible because correspondence to objective metrics improved up to 10 clips without saturating. Condition-level aggregation, confidence-interval reporting, and, when bimodality is present, stratification by positive versus negative preference mode are all supported by the reported analyses (Suárez et al., 2023, Naderi et al., 2020).

A further methodological caution concerns historical baselines. In the speech crowdsourcing reproducibility study, comparison with historical CCR laboratory results from ITU-T Supplement P.23 produced very low correlation, with Pearson AA6, and the historical data were argued to underweight background-noise effects. The stated conclusion was that those historical lab CCR results should not be used as a present-day ground truth for noise-rich conditions. This controversy underscores a broader point: CCR outcomes are sensitive not only to scaling and presentation, but also to playback conditions, listener expectations, and the perceptual ecology of the test era (Naderi et al., 2021).

CCR is therefore best regarded as a high-resolution comparative instrument whose validity depends on rigorous control of presentation, screening, aggregation, and interpretation. When those elements are handled carefully, the method yields reproducible condition-level judgments, captures both improvements and degradations, and exposes heterogeneity in human trade-offs that absolute metrics or single-stimulus ratings may obscure.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Comparison Category Rating (CCR).