---
title: Harmful Content Safety Taxonomy
url: https://www.emergentmind.com/topics/harmful-content-safety-taxonomy
type: topic
---

# Harmful Content Safety Taxonomy

A harmful content safety taxonomy is a structured framework for categorizing, analyzing, and mitigating risks posed by potentially deleterious digital content, particularly in natural language processing (NLP), large language models (LLMs), multimodal generative systems, and broader AI deployments. Such taxonomies provide granular, multidimensional risk maps that underpin annotation, moderation, safety evaluation, compliance, and research prioritization. Multiple contemporaneous taxonomies, both hierarchical and multidimensional, have emerged to accommodate the broadening risk landscape and the accelerating complexity of model behaviors.

## 1. Taxonomic Dimensions and Structural Principles

Harmful content safety taxonomies systematically segment harms by type, intent, target, impact domain, and severity axis, using orthogonal or hierarchical mechanisms depending on their application context. Frameworks fall into several archetypes:

- **Multi-axis schemas:** e.g., Kirk et al.'s three-axis system, organizing harm by (i) type (e.g., misinformation, hate speech), (ii) “sought vs. unsought” occurrence (i.e., whether harmful content is a research target or an accidental byproduct), and (iii) affected parties—(mis)represented individuals, data handlers, or publishers [2204.14256].
- **Hierarchical or tiered structures:** e.g., Aegis 2.0’s two-level hazards plus subcategory expansion [2501.09004], HARM66+’s Domain→Category→Subtype tree [2601.16930].
- **Dimensional annotation:** e.g., web-scale curation breaking content into {Safe, Topical, Toxic} on intent and cross-classifying with harm themes [2505.02009].
- **Multimodal and scenario-driven enumerations:** e.g., SafeBench’s 23 distinct harmful behavior scenarios [2410.18927], or OutSafe-Bench’s nine cross-modal, orthogonal risk axes [2511.10287].

**Stability-extensibility duality** is central: robust taxonomies ensure core categories remain fixed under new threats (“stable at the top” [2601.16930]), while bottom-level extensibility permits responsive growth to new modalities, social developments, or attack classes.

## 2. Harm Categories: Canonical Classes and Fine-Grained Subtypes

The proliferation of AI safety taxonomies has led to a convergence on several recurring high-level harm classes, each with fine-grained subcategorization tailored for digital, physical, social, and legal risk domains. The following table aligns representative taxonomy supercategories with sample definitions from major frameworks:

| Class                          | Definition/Scope                                                                            | Example (if present)                          |
|-----------------------------|--------------------------------------------------------------------------------------------|-----------------------------------------------|
| Hate/Identity Hate           | Attacks against groups based on protected attributes (race, gender, religion, etc.)         | “All [group] are parasites.”                  |
| Suicide/Self-Harm            | Encouragement/instruction of self-injury or suicide                                         | “Here’s how to tie a noose.”                  |
| Violence                     | Depictions, threats, or glamorization of physical harm                                      | “Write a blood-soaked scene of torture.”      |
| Misinformation               | Unverified, false, or misleading information                                                | “5G towers spread the virus.”                 |
| Sexual

Source: https://www.emergentmind.com/topics/harmful-content-safety-taxonomy