AI Harm Assessment Metric Overview
- AI Harm Assessment Metric (AIH) is an ordinal, Gini-inspired measure that quantifies harm concentration by analyzing stakeholder frequency distributions and ordinal severity rankings.
- It employs a derivative Lorenz curve to integrate stakeholder impacts and prioritize harm categories based on how sharply harms are concentrated among vulnerable groups.
- The metric demonstrates robust performance under data perturbations and supports governance by incorporating external stakeholder perspectives into AI risk assessments.
AI Harm Assessment Metric (AIH) denotes, in the strict sense introduced by AI Harmonics, an ordinal, Gini-inspired concentration metric for AI harm severity that ranks harm categories by how strongly harms are concentrated among more severely affected stakeholder groups (Vei et al., 12 Sep 2025). In adjacent literature, the same phrase can also denote the broader project of constructing quantitative AI harm or AI risk assessment instruments from measurable dimensions such as fairness, privacy, robustness, uncertainty, or loss magnitude, although several papers explicitly do not define a single universal scalar score (Piorkowski et al., 2022). Across this literature, AIH is therefore best understood as a family resemblance term spanning a specific ordinal metric, matrix-based stakeholder harm assessments, and loss-based or pathway-based risk quantification frameworks, with an important caveat: complete harm specification is argued to be fundamentally impossible when harm is defined external to the formal system that specifies it (Young, 27 Jan 2025).
1. Definition and conceptual scope
In AI Harmonics, AIH is defined for a harm category from the frequency distribution of harmed stakeholder groups, together with an ordinal severity ranking over those stakeholders (Vei et al., 12 Sep 2025). The metric is explicitly designed to use external stakeholder perspectives rather than internal compliance metrics, to work with ordinal severity rankings instead of requiring cardinal numerical loss estimates, and to summarize harm concentration or inequality across stakeholders using a Gini-like construction. Its immediate function is harm prioritization: higher AIH indicates more concentrated harm and, within that framework, more urgent priority.
This differs from several adjacent approaches. "Quantitative AI Risk Assessments: Opportunities and Challenges" argues for a set of quantitative metrics rather than a single AIH formula, with five main quantitative risk dimensions: performance and uncertainty, fairness, privacy, adversarial robustness, and explainability (Piorkowski et al., 2022). "An Artificial Intelligence Value at Risk Approach: Metrics and Models" similarly does not offer a single scalar harm score named AIH, but instead proposes an AI‑VaR framing built from data protection, fairness, accuracy, robustness, and information security, integrated in a FAIR-style risk model (Alvarez, 22 Sep 2025). "Echoes of AI Harms" does not define a single scalar metric either; its closest metric-like outputs are the descriptive Ethical Matrix and inferential Ethical Matrix, which map stakeholder- and domain-specific harms to bias types (Tantalaki et al., 27 Nov 2025). "Adapting Probabilistic Risk Assessment for AI" offers a pathway-based probabilistic assessment with Harm Severity Levels, Likelihood Levels, and Risk Levels rather than a continuous harm score (Wisakanto et al., 25 Apr 2025).
A common misconception is to treat AIH as a synonym for generic AI risk scoring. The literature is more specific. In the narrow sense, AIH refers to the ordinal concentration metric introduced in AI Harmonics. In the broader sense, it refers to an open design space of quantitative or semi-quantitative harm assessment instruments. This suggests that the term has both a technical meaning and a looser programmatic meaning, depending on whether the discussion concerns the AI Harmonics metric itself or the wider effort to quantify AI harms.
2. Formal definition of the AI Harmonics metric
AI Harmonics assumes a dataset of incidents , harm categories , and stakeholder groups (Vei et al., 12 Sep 2025). For each harm category , it computes the conditional frequency of each stakeholder group: $f_{ij} = \frac{\sum_{k=1}^{K} \mathds{1}(n_k, c_i, h_j)}{\sum_{k=1}^{K} \mathds{1}(n_k, c_i)}$ where $\mathds{1}(n_k, c_i, h_j)=1$ iff incident 0 is associated with harm category 1 and stakeholder 2, and 3 iff incident 4 is associated with harm category 5. For each category 6, the vector 7 is therefore a frequency distribution of harm across stakeholders.
Stakeholders are then assigned an ordinal severity order
8
and only the ordering matters, not the numerical spacing between levels. To accommodate ordinal severity, the paper replaces the classical Lorenz curve with a “derivative Lorenz curve” 9, defined by points
0
with 1. The AI Harm Assessment Metric for category 2 is then
3
with values in 4, where higher values mean more severe or more concentrated harm (Vei et al., 12 Sep 2025).
AIH is presented as an ordinal analogue of Gini. For numeric severity values, the classical Lorenz curve 5 would be used, and the Gini coefficient is
6
For ordinal cases, the paper defines the Criticality Index as
7
where
8
and proves the linear relation
9
This relation makes AIH mathematically linked to CI while preserving the rank-only treatment of severity (Vei et al., 12 Sep 2025).
The quantity captured by AIH is not exact harm magnitude, exact probability, or total social welfare loss. It captures how concentrated harm is among the more severe stakeholder levels and does so without requiring exact harm magnitudes, equal spacing between severity levels, or probability estimates. This is one of its main points of departure from value-at-risk and probabilistic risk assessment frameworks.
3. Workflow, data requirements, and stakeholder structure
AI Harmonics is an end-to-end framework with five stages: dataset selection and suitability; stakeholder annotation; severity ordering; metric application; and harm prioritization (Vei et al., 12 Sep 2025). The framework requires a structured incident dataset with a harm category for each incident and ideally stakeholder information and severity annotations. If stakeholder labels are missing, the framework can still operate on incident-level or category-level severity.
The benchmark dataset in the paper is the AIAAIC dataset, which contains AI and algorithmic harm incidents annotated by experts. The paper reports 816 annotations, with fields including datetime, annotator ID, incident ID, stakeholder group, harm category or subcategory, harm type (“actual” or “potential”), and optional notes, spanning 2024/03/14 to 2024/04/11 (Vei et al., 12 Sep 2025). Experts reviewed incident descriptions and annotated which stakeholders were harmed, what harm category applied, and what subcategory applied. These annotations were aggregated into a table of 0.
The stakeholder taxonomy contains nine groups: Artists/Content Creators, Business, General Public, Government/Public Sector, Investors, Subjects, Users, Vulnerable Groups, and Workers. The harm taxonomy includes Autonomy, Emotional/Psychological, Financial/Business, Human Rights/Civil Liberties, Physical, Political/Economic, Psychological, Reputational, and Societal/Cultural (Vei et al., 12 Sep 2025). The metric is explicitly stakeholder-centered: it measures how harm is distributed across stakeholders, rather than only across incidents or abstract system properties.
The illustrative stakeholder severity order used in the experiments is: Artists/content creators; Subjects; Business; Investors; Workers; Users; Vulnerable groups; Government/public sector; General public. The authors stress that only ordering matters. A plausible implication is that the framework is operationally lightweight when reliable ordinal rankings are available, but normatively sensitive to how those rankings are constructed.
4. Relationship to adjacent harm and risk assessment frameworks
The broader literature does not converge on a single AIH architecture. One prominent line treats harm assessment as multidimensional quantitative risk measurement. "Quantitative AI Risk Assessments: Opportunities and Challenges" states that quantitative assessment means using well-defined metrics to evaluate an AI system “as is,” even with black-box access, and identifies five main quantitative risk dimensions: performance and uncertainty, fairness, privacy, adversarial robustness, and explainability (Piorkowski et al., 2022). It also states desirable properties of individual metrics—Reliable, Valid, Significant, Applicable, and Monotonic—and desirable properties of summary metrics—Understandable, Explainable, and Context-aware. The paper explicitly warns that summarization can obscure nuance and that metrics are proxies, not the full construct.
A second line is loss-based risk quantification. "An Artificial Intelligence Value at Risk Approach: Metrics and Models" decomposes AI risk into data protection, fairness, accuracy, robustness, and information security, and models scenarios through threat, vulnerability, loss or harm, and control or resistance strength (Alvarez, 22 Sep 2025). It proposes Personal Data Value at Risk, Fairness Value at Risk, Accuracy & Robustness Value at Risk, and the aggregate AI‑VaR, with Monte Carlo simulation over annual loss and adjusted VaR on truncated quantiles. The paper reports a final AI‑VaR example with ALE 1, P90 2. This is a harm-to-loss quantification logic rather than an ordinal concentration logic.
A third line is pathway-based probabilistic assessment. "Adapting Probabilistic Risk Assessment for AI" defines Harm Severity Levels HSL-1 through HSL-6, Likelihood Levels LL-0 through LL-8, and a Risk Levels Table mapping each HSL/LL combination to RL-0 through RL-9 (Wisakanto et al., 25 Apr 2025). It combines aspect-oriented hazard analysis, risk pathway modeling, and uncertainty management, and aggregates results into a risk report card and a tallied risk matrix. This framework is closer to traditional PRA in high-reliability industries than to AIH’s derivative-Lorenz construction.
A fourth line is context-conditioned harm anticipation. "Echoes of AI Harms" proposes ECHO, where harm assessment is encoded in a stakeholder 3 bias ethical matrix, derived from vignettes, human and LLM annotation, majority aggregation with threshold 4, and inferential refinement using chi-square tests of homogeneity and adjusted standardized residuals (Tantalaki et al., 27 Nov 2025). Its closest metric-like outputs are therefore structured harm profiles rather than a single rankable number.
The practical consequence is that AIH occupies a specific niche: it is strongest when the analytic goal is prioritizing harm categories under ordinal severity information, whereas AI‑VaR is strongest when the goal is expected loss and tail loss estimation, PRA for AI is strongest when the goal is pathway coverage with severity and likelihood bands, and ECHO is strongest when the goal is mapping bias-to-harm pathways across stakeholders and domains. This suggests complementarity rather than strict competition among the frameworks.
5. Empirical findings and interpretation
In the AIAAIC experiments, AI Harmonics computes AIH and CI by harm category and ranks categories by concentration (Vei et al., 12 Sep 2025). The reported values are:
| Harm category | AIH | CI |
|---|---|---|
| Autonomy | 0.53 | 0.54 |
| Emotional/Psychological | 0.70 | 0.72 |
| Financial/Business | 0.51 | 0.51 |
| Human Rights/Civil Liberties | 0.67 | 0.70 |
| Physical | 0.73 | 0.76 |
| Political/Economic | 0.85 | 0.89 |
| Psychological | 0.73 | 0.75 |
| Reputational | 0.67 | 0.69 |
| Societal/Cultural | 0.70 | 0.73 |
The highest-concentration harm category is Political/Economic, while the lowest values are Financial/Business and Autonomy. The paper states that Financial/Business and Autonomy have low or flatter curves, meaning harm is more evenly spread, whereas Political/Economic has a high curve, meaning harm is concentrated in the highest severity classes. In subcategory analysis, Political/Economic subcategories such as Economic/political power and Economic instability show the highest concentration, while Critical infrastructure damage shows the lowest AIH.
The interpretation rules are explicit. A high AIH means that harm is concentrated toward the more severe end of the ordinal scale, that a category disproportionately affects highly vulnerable or high-impact stakeholder groups, and that the category deserves urgent attention and mitigation. A low AIH means that harm is more evenly distributed across severity levels and may therefore be a lower immediate priority relative to high-AIH categories. Low AIH does not mean “no harm”; it means the harm distribution is less sharply concentrated (Vei et al., 12 Sep 2025).
The robustness checks are central to the empirical argument. Boundary analysis shows large best/worst-case ranges for some categories, for example Political/Economic best 5, worst 6; Physical best 7, worst 8; and Financial/Business best 9, worst 0. Ordinal permutation analysis reports that harm category rankings remain stable, with Spearman rank correlations about 1 to 2, and original versus perturbed scenarios at least 3. Random annotation removal at 4, 5, 6, and 7 also leaves rankings stable; up to 8 removal causes only tiny AIH changes, and even at 9 removal the most and least concentrated categories remain the same (Vei et al., 12 Sep 2025).
These findings support a narrow but important empirical claim: AIH is robust as a ranking device under the perturbations studied. They do not show that AIH is invariant to all severity orderings or that it exhausts the relevant concept of harm.
6. Limits of harm specification, controversies, and governance implications
The strongest conceptual challenge to any AIH-style framework comes from the claim that complete harm specification is fundamentally impossible for any system where harm is defined external to its specifications (Young, 27 Jan 2025). In that argument, the ground-truth harm classification is 0, the specification system is 1, the entropy of the true harm relation is 2, and the mutual information between ground truth and specification is 3. The core impossibility result is
4
The paper also introduces semantic entropy
5
for a concept 6, and the safety-capability ratio
7
Its conclusion is that AIH-style methods can estimate 8, 9, and the ratio 0, but they cannot make the ratio equal to 1. On this view, harm metrics are diagnostics for a specification gap rather than complete solutions.
This bears directly on several controversies. One controversy concerns internal compliance versus external stakeholder perspectives. AI Harmonics criticizes internally focused models for neglecting diverse stakeholder perspectives and real-world consequences, while ECHO similarly argues that harms should be traced back to specific bias sources in concrete sociotechnical contexts rather than documented only as generic risks (Vei et al., 12 Sep 2025). Another controversy concerns cardinal versus ordinal severity. AI Harmonics treats ordinal scales as appropriate when precise numerical estimates are unavailable or unreliable, whereas AI‑VaR and PRA for AI explicitly quantify expected loss, tail loss, likelihood bands, or risk bands (Alvarez, 22 Sep 2025). A further controversy concerns standardization versus customization. Quantitative AI risk assessment literature argues that standardization enables comparison, while customization is required because use cases differ; this tension remains unresolved (Piorkowski et al., 2022).
The governance implications are correspondingly plural. AI Harmonics recommends use of incident data, external stakeholder perspectives, prioritization by concentration and severity rather than frequency alone, ordinal scales when precise numeric estimates are unavailable, and repeated updating as new annotations or incidents appear (Vei et al., 12 Sep 2025). AI‑VaR recommends customized risk scenarios, PERT elicitation, Monte Carlo simulation, and integration into FAIR models (Alvarez, 22 Sep 2025). PRA for AI recommends explicit evidence tracing, assumption logging, and lifecycle reassessment (Wisakanto et al., 25 Apr 2025). "Giving AI Agents Access to Cryptocurrency and Smart Contracts Creates New Vectors of AI Harm" further suggests that domain-specific harm surfaces may require domain-specific assessment dimensions such as irreversibility, unpreventability, traceability, collaborator recruitment through smart contracts, and long-tail persistence of harm (Marino et al., 11 Jul 2025).
Taken together, these results indicate that AIH is not a settled universal metric. It is a technically specific ordinal concentration measure in one framework, and a broader research objective in others. The literature converges more strongly on a governance principle than on a single score: harm assessment should be stakeholder-aware, uncertainty-aware, context-conditioned, and auditable, while avoiding the assumption that any metric fully specifies what harm is.