Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
102 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Placing M-Phasis on the Plurality of Hate: A Feature-Based Corpus of Hate Online (2204.13400v1)

Published 28 Apr 2022 in cs.CL

Abstract: Even though hate speech (HS) online has been an important object of research in the last decade, most HS-related corpora over-simplify the phenomenon of hate by attempting to label user comments as "hate" or "neutral". This ignores the complex and subjective nature of HS, which limits the real-life applicability of classifiers trained on these corpora. In this study, we present the M-Phasis corpus, a corpus of ~9k German and French user comments collected from migration-related news articles. It goes beyond the "hate"-"neutral" dichotomy and is instead annotated with 23 features, which in combination become descriptors of various types of speech, ranging from critical comments to implicit and explicit expressions of hate. The annotations are performed by 4 native speakers per language and achieve high (0.77 <= k <= 1) inter-annotator agreements. Besides describing the corpus creation and presenting insights from a content, error and domain analysis, we explore its data characteristics by training several classification baselines.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (9)
  1. Dana Ruiter (11 papers)
  2. Liane Reiners (2 papers)
  3. Ashwin Geet D'Sa (4 papers)
  4. Thomas Kleinbauer (7 papers)
  5. Dominique Fohr (13 papers)
  6. Irina Illina (18 papers)
  7. Dietrich Klakow (114 papers)
  8. Christian Schemer (1 paper)
  9. Angeliki Monnier (1 paper)
Citations (2)

Summary

We haven't generated a summary for this paper yet.