Papers
Topics
Authors
Recent
Search
2000 character limit reached

Majority Vote Silences Minority Values: Annotator Disagreement at the Hate/Offensive Boundary in HateXplain

Published 27 Jun 2026 in cs.CL and cs.AI | (2606.28772v1)

Abstract: Hate speech annotation pipelines routinely collapse annotator disagreement into majority vote labels before training. We show that this aggregation is not neutral: 42.6% of all annotator disagreement in HateXplain concentrates specifically at the hate/offensive boundary, a pattern consistent with annotators applying different thresholds for where hate begins (chi-squared = 135.199, df = 2, p < 0.0001). Both a hard-label BERT model (Model A) and a soft-label model (Model B) drop 22 percentage points in accuracy from agreed posts (~80%) to disagreement posts (~58%), confirmed at p < 0.0001. A per-annotator multi-head model (Model C) widens this gap further to 28 points while collapsing offensive disagreement accuracy to 0.245. Critically, Model A expresses significantly higher confidence on boundary case errors than Model C (0.710 vs. 0.495, p < 0.0001), meaning standard evaluation metrics will not detect the failure. Three downstream interventions of increasing sophistication all fail to recover boundary accuracy. We argue the problem is structural. Majority vote presents a contested judgment as ground truth, and models inherit that false certainty. The intervention must be upstream in annotation design.

Summary

  • The paper quantifies that 42.6% of annotator disagreement focuses on the hate/offensive boundary, highlighting majority-vote aggregation’s masking of minority values.
  • The paper shows a consistent 22–28 point accuracy gap between full agreement and disagreement cases, with per-annotator models lowering offensive error confidence but not improving boundary resolution.
  • The paper underscores that advanced modeling techniques fail to recover minority perspectives, calling for revised annotation schemas and pluralistic alignment in dataset design.

Majority Vote Aggregation and Disagreement at the Hate/Offensive Boundary in HateXplain

Overview and Motivation

This paper presents a rigorous analysis of annotation disagreement at the hate/offensive speech boundary in the HateXplain corpus, with a focus on the consequences of majority-vote label aggregation for model calibration and pluralistic value representation. The investigation is motivated by the observation that decisions at the hate/offensive boundary are highly subjective—reflecting annotator-specific thresholds rather than empirical facts—yet standard machine learning pipelines flatten this disagreement, presenting the majority's viewpoint as "ground truth". The study not only quantifies the structural locus of annotator disagreement within HateXplain but also empirically evaluates whether downstream modeling interventions can remedy the informational losses induced by majority aggregation.

Disagreement Structure and Localization

A principal finding is the non-uniform distribution of annotator disagreement: 48.7% of HateXplain posts exhibit disagreement, with 42.6% of all disagreement focusing specifically at the hate/offensive boundary—a statistically significant concentration (χ2=135.199\chi^2 = 135.199, df=2df = 2, p<0.0001p < 0.0001). Disagreement is particularly acute for offensiveness judgments: offensive posts disagree at 67.9%, compared with 50.5% for hatespeech and 35.5% for normal content.

The distribution of soft label probabilities reveals discrete mass concentrated on the hate/offensive axis, indicating that disagreement reflects genuine value thresholds rather than stochastic labeling noise.

Figure 1

Figure 1: Annotator disagreement rates in HateXplain, label distribution by agreement, and soft label mass at the hate/offensive boundary.

Model-Based Analysis: Agreement Gap and Calibration Shortcomings

Three BERT-base models were trained and compared using majority-vote (Model A), soft label (Model B), and per-annotator head (Model C) approaches. Across all approaches, there is a consistent accuracy gap of approximately 22–28 points between examples with full annotator agreement (∼\sim80% accuracy) and those exhibiting disagreement (∼\sim55–58% accuracy). This gap is robust (all p<0.0001p < 0.0001) and widens under the strongest per-annotator baseline.

Notably, per-annotator modeling (Model C) further collapses accuracy on posts with offensive disagreement (to 0.245), despite improved calibration—Model C lowers its mean prediction confidence on boundary errors from 0.710 (Model A) to 0.495. However, this greater epistemic humility does not translate into more accurate boundary resolution.

Figure 2

Figure 2: Accuracy by agreement level, disagreement category accuracy, and boundary error confidence across modeling approaches.

Majority alignment analysis further demonstrates that models remain strongly majority-aligned in their predictions (∼\sim70% on disagreement posts), and the gap between agreement and disagreement cases is unaffected or exacerbated by more complex modeling.

Figure 3

Figure 3: Model alignment with majority/minority annotators and the persistent/widening gap with more sophisticated models.

Annotation Error Versus Value Disagreement

To directly interrogate whether boundary disagreements are attributable to annotation error or genuine subjective thresholding, the study leverages token-level rationale annotations. Pairs of annotators who chose different labels for the same post overlap in highlighted tokens as frequently (or more so) than those who agreed—73.1% of such pairs have Jaccard overlap ≥0.3\geq 0.3, and more than half reach ≥0.5\geq 0.5. Crucially, a subset of annotator pairs highlight identical tokens yet diverge in label assignment, providing a clear signature of value-based (threshold) disagreement rather than evidence divergence.

Figure 4

Figure 4: Jaccard overlap distributions of highlighted tokens among annotator pairs: high overlap among both agreeing and disagreeing pairs argues against noise as the dominant source of disagreement.

Implications for Pluralistic Alignment and Dataset Design

The findings have broad implications for the design of both datasets and machine learning systems in sociotechnical domains where value judgments are central. Majority-vote aggregation not only masks persistent, location-specific disagreement but also leads to overconfident models that inherit and ossify the perspective of the majority—a form of epistemic closure that systematically suppresses minority annotator values. Downstream modeling interventions—even those explicitly designed to attend to annotation heterogeneity—do not recover lost information or minority perspectives once majority aggregation has occurred, highlighting the fundamentally upstream nature of the issue.

The evidence from this analysis directly supports calls for pluralistic alignment in AI, as represented by work on jury learning and graduated harm scales. Three specific recommendations are advanced:

  • Utilize graduated harm scales rather than categorical hate/offensive boundaries, to more faithfully capture value-contingent thresholds and reduce boundary-specific disagreement.
  • Surface model uncertainty on high-disagreement inputs to enable human-in-the-loop review, rather than relying solely on model predictions.
  • Ensure that annotation pools represent target communities, especially those systematically impacted by moderation policies.

The paper also notes important ethical considerations regarding annotation work on sensitive content, emphasizing the need for proper compensation, opt-in participation, and mental-health safeguards for annotators from marginalized groups.

Limitations and Future Work

While the study provides robust evidence within the HateXplain context, it is limited by the lack of annotator demographic metadata (precluding analysis of the sociocultural roots of threshold variation) and by the use of BERT-base exclusively. The proposed hypothesis—that value-laden aggregation failures generalize to other subjective ML pipelines (e.g., RLHF reward modeling)—remains to be tested in broader domains and with richer annotations, such as those available in the DICES dataset.

Conclusion

This work demonstrates that majority-vote aggregation in hate speech annotation elides substantial, targeted disagreement at critical value boundaries, leading to models that are not only inaccurate on contested inputs but also inappropriately confident. Downstream technical remedies fail to address the loss, underscoring the need for upstream interventions in annotation schema, pool composition, and label design to ensure the proper representation of plural and minority value perspectives in AI models.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.