AI Harmonics: Stakeholder-Centric Risk Metric
- AI Harmonics is a harm-centric, stakeholder-adaptive risk framework that uses an ordinal AIH metric to rank harms based on empirical incident data.
- The framework emphasizes external stakeholder impacts by re-ranking harms dynamically, helping prioritize regulatory and mitigation efforts.
- Empirical analysis shows high concentration in Political & Economic harms, with robust stability even under significant data perturbations.
AI Harmonics is a stakeholder-centric, harm-severity adaptive AI risk assessment framework built around the AI Harm assessment metric, AIH, an ordinal measure designed to prioritize harms from empirical incident data without requiring precise utilities or probabilities. It starts from expert-annotated AI incident repositories, represents who is harmed, how harms are categorized, and how severely they should be ranked on an ordinal scale, and then ranks harm categories by the concentration of incidents at high severity. In its 2025 formulation, the framework is presented as a data-driven, external-viewpoint alternative to provider-centric and compliance-driven risk assessment, with empirical application to AIAAIC annotations showing especially strong concentration for Political & Economic harms and substantial concentration for Physical and Psychological harms (Vei et al., 12 Sep 2025).
1. Definition and conceptual orientation
AI Harmonics is described as a “harm-centric and stakeholder-adaptive AI risk assessment framework that embeds adopters’ views about AI harms captured over ordinal scales”. Its central claim is methodological rather than merely taxonomic: AI harms should be assessed from real-world incidents and stakeholder impacts, not only from internal controls, provider liabilities, or compliance checklists. This positions the framework against approaches that emphasize internal audits and abstract risk classes while under-representing external stakeholders, affected communities, and experienced harms (Vei et al., 12 Sep 2025).
The framework is explicitly human-centric. It starts from AI incidents understood as “physical, psychological, or otherwise consequential damage” to individuals, groups, organizations, society, or the environment. Severity is framed as the “magnitude of negative impacts on assets, operations, individuals, organizations, the Nation, or society more broadly.” The adaptive aspect lies in re-ranking harms as new incidents, annotations, or stakeholder priorities change, rather than fixing systems into static risk tiers. This suggests that AI Harmonics is intended as a continuously updateable prioritization layer rather than a one-time classification exercise.
A recurring misconception is to treat AI Harmonics as a replacement for all other AI governance machinery. The framework is narrower and more specific: it provides a stakeholder-aware, ordinal, incident-grounded way to rank harms by severity concentration. It does not eliminate the need for legal analysis, system evaluation, or domain-specific oversight. Another common misunderstanding is to read AIH as a raw incident-frequency score. The framework instead computes conditional frequencies within each harm category and asks how much of that category is concentrated at higher severity ranks.
2. Stakeholders, harm taxonomy, and severity ordering
The framework models harms through three linked elements: stakeholder groups, harm categories or subcategories, and ordinal severity levels. Stakeholders are discrete groups such as Artists/Content Creators, Business, General Public, Government/Public Sector, Investors, Subjects, Users, Vulnerable Groups, and Workers. Multiple stakeholder annotations per incident are allowed, so a single incident may affect several groups at once (Vei et al., 12 Sep 2025).
Severity is operationalized ordinally rather than numerically. The illustrative stakeholder ranking given in the framework is as follows.
| Stakeholder group | Severity rank |
|---|---|
| Artists/Content creators | 1 |
| Subjects | 2 |
| Business | 3 |
| Investors | 4 |
| Workers | 5 |
| Users | 6 |
| Vulnerable groups | 7 |
| Government/public sector | 8 |
| General public | 9 |
These values are purely ordinal: only the order matters, not the distances between ranks. This design choice addresses a central problem in AI-risk practice identified by the authors: severity labels are often ambiguous, inconsistent, or noisy, and numeric scoring such as “7/10 severity” is arbitrary and not comparable across sources.
AI Harmonics uses the AIAAIC harm taxonomy. High-level categories include Autonomy; Physical; Psychological / Emotional psychological; Reputational; Financial & Business; Human rights & Civil liberties; Societal & Cultural; and Political & Economic. The experimental analysis reports nine high-level categories: Autonomy, Emotional & Psychological, Financial & Business, Human Rights & Civil Liberties, Physical, Political & Economic, Psychological, Reputational, and Societal & Cultural. Subcategories include, for example, loss of life, bodily injury, property damage, coercion, anxiety/distress, defamation/libel/slander, financial/earnings loss, discrimination, public service deterioration, political instability, institutional trust loss, critical infrastructure damage, political manipulation, economic instability, and electoral interference.
This structure makes the framework simultaneously normative and empirical. It is normative because the severity ordering reflects value judgments about whose harms should count as more severe in allocation decisions. It is empirical because the ordering is applied to annotated incident data rather than abstract hypothetical scenarios.
3. Mathematical formulation of the AIH metric
For a harm category and stakeholder groups , , ordered by severity
AI Harmonics first computes conditional frequencies
$f_{ij} = \frac{\sum_{k=1}^{K}\mathds{1}(n_k, c_i, h_j)} {\sum_{k=1}^{K}\mathds{1}(n_k, c_i)},$
where is the number of incidents, $\mathds{1}(n_k, c_i, h_j)=1$ if incident is tagged with category and stakeholder , and 0 if incident 1 has category 2. Thus 3 is the share of category-4 incidents affecting stakeholder 5 (Vei et al., 12 Sep 2025).
The framework then defines a stepwise “derivative” Lorenz curve by points
6
The AIH score for category 7 is the area under that curve: 8 Its range is 9. Higher values indicate that harms in a category are concentrated among higher-severity stakeholders; lower values indicate more even spread across severity levels.
The formalism is deliberately ordinal. In the numeric case, with meaningful severity magnitudes 0, one may instead use the classical Lorenz curve
1
and the standard Gini index
2
AI Harmonics rejects that move when severity intervals are not meaningful. The framework therefore treats AIH as a “pseudo-Gini” for ordinal data.
The paper also establishes a linear relation with the Criticality Index: 3 and
4
Because AIH is a rescaled 5, both metrics yield the same ordering of categories. This provides continuity with ordinal cyber-risk modeling while retaining a formulation tailored to AI incidents and stakeholders.
4. Workflow and empirical data model
The operational pipeline has five main stages: dataset selection and suitability; stakeholder annotation; severity ordering; distribution computation; and metric application with harm prioritization. The framework is dataset-agnostic and is designed to work with AIAAIC, AI Incident Tracker, OECD AI Incidents Monitor, and related repositories, provided that harm categories are available and stakeholder or severity annotations can be supplied through experts or LLM assistance if necessary (Vei et al., 12 Sep 2025).
In the main empirical demonstration, the source dataset is AIAAIC over the period 2024-03-14 to 2024-04-11. The data comprise 816 expert annotations in a many-to-one relation with incidents. Each annotation includes datetime, annotator ID, incident ID, stakeholder group(s), harm category and subcategory, harm type (“actual” or “potential”), and notes. For metric computation, these annotations are condensed into frequency tables of the form 6.
This aggregation is methodologically significant. AI Harmonics does not require exact causal estimates, calibrated probabilities, or utility functions. It needs structured annotations that support relative ranking. If stakeholder labels are unavailable, the framework can instead treat incidents themselves or harm types as the units to be ranked by severity. A plausible implication is that the method is especially suited to settings where qualitative evidence is abundant but precise quantitative risk models are not.
The framework’s design also embeds social context. One incident can be tagged with multiple stakeholder groups and multiple harm categories. That allows it to represent harms that are distributed unevenly across populations, including cases where users, vulnerable groups, or the general public absorb most of the burden while providers remain comparatively insulated from direct consequences.
5. Empirical findings and robustness properties
The empirical results reported for AIAAIC show uneven harm distributions across categories. The following values are given for AIH and 7.
| Harm Category | AIH | CI |
|---|---|---|
| Autonomy | 0.53 | 0.54 |
| Emotional & psychological | 0.70 | 0.72 |
| Financial & business | 0.51 | 0.51 |
| Human rights & civil liberties | 0.67 | 0.70 |
| Physical | 0.73 | 0.76 |
| Political & economic | 0.85 | 0.89 |
| Psychological | 0.73 | 0.75 |
| Reputational | 0.67 | 0.69 |
| Societal & cultural | 0.70 | 0.73 |
The principal empirical conclusion is that Political & Economic harms exhibit the highest harm-severity concentration, with Physical and Psychological harms also highly concentrated. Political & Economic harms include institutional trust loss, electoral interference, political manipulation, and economic/political power. Physical harms include loss of life, bodily injury, and health deterioration. The interpretation offered is that these categories warrant urgent mitigation because they are concentrated among higher-severity stakeholders and involve severe real-world consequences (Vei et al., 12 Sep 2025).
The underlying frequency structure supports that reading. The heatmap summary reports, for example, Human rights & civil liberties with 40 incidents for Vulnerable groups and 39 for Users; Autonomy with 29 incidents for Users and 15 for Artists/Content creators; and Physical with 22 incidents for Vulnerable groups and 14 for Users. Investors and Subjects show zero or near-zero counts in several categories. This indicates disproportionate impact on Users and Vulnerable groups across multiple harm domains.
Robustness analysis is a major part of the framework’s empirical argument. Under random removal of annotations from 10% to 80%, category rankings remain stable at the extremes: Political & Economic remains the highest category, while Financial & Business and Autonomy remain the lowest. Under random permutations of the stakeholder severity ordering, mean AIH per category remains stable, and Spearman rank correlation across scenarios satisfies 8. Within Political & Economic subcategories, Economic/political power and Economic instability have the highest median AIH, Critical infrastructure damage the lowest AIH, and Electoral interference a rather stable AIH with small variance across scenarios. This suggests that the framework is robust both to incomplete data and to uncertainty in the exact ordinal ordering.
6. Governance uses, limitations, and terminological ambiguity
For policymakers and regulators, AI Harmonics is proposed as a way to prioritize regulatory focus, identify uneven harm distributions, allocate oversight resources, and support dynamic risk governance. For organizations, developers, deployers, and auditors, it can be used for internal harm mapping, mitigation prioritization, scenario and sensitivity analysis, and integration with risk matrices, impact assessments, ISO/NIST risk frameworks, and EU AI Act risk tiers. The framework also includes open-source code, an interactive web app, and a pre-processed AIAAIC benchmark dataset (Vei et al., 12 Sep 2025).
Its limitations are explicit. Results depend on the coverage and representativeness of incident repositories, which may be biased toward Western or English-language reporting and certain sectors. Annotation of harm categories, stakeholders, and severity orderings is subjective and affected by epistemic uncertainty. Rare catastrophic harms are difficult to capture because AIH is a concentration metric rather than a model of tail-risk probability. The framework does not explicitly model likelihood of occurrence in a deployment context. Severity must be meaningfully rank-orderable, and the chosen stakeholders, categories, and severity levels must reflect the normative priorities of the community applying the method. This suggests that AIH is best read as a comparative prioritization signal within an incident corpus, not as a complete deployment-specific risk estimate.
The name also admits a technical ambiguity. In a distinct power-systems context, research on active power filter control for non-linear non-stationary currents used empirical mode decomposition as an adaptive, data-driven method for separating harmonics and disturbances before applying modified 9-0 theory, and described that architecture as conceptually aligned with broader “AI-enabled” harmonic analysis and mitigation (Tai et al., 2012). That usage concerns power-quality control, not AI governance. The contemporary term AI Harmonics, however, denotes the human-centric, stakeholder-aware risk assessment framework organized around the AIH metric.