---
title: Moral Graph Elicitation
url: https://www.emergentmind.com/topics/moral-graph-elicitation
type: topic
---

# Moral Graph Elicitation

Moral graph elicitation refers to a set of rigorous, graph-structured methodologies for extracting, representing, and reconciling moral values—whether from human populations, news text, or AI systems—using explicit graph schemas and computational workflows. These approaches operationalize values and their relationships as nodes and edges, facilitating subsequent alignment, interpretability, or utility analysis. Moral graph elicitation underpins both the design of scalable value alignment protocols for AI and the systematic study of moral event structure in natural language.

## 1. Formal Definitions of Moral Graphs

A moral graph is a structured, directed representation encoding values, agents, events, and/or comparative value preferences. Distinct formalizations are found in the literature.

### Moral Graph (Human Value Elicitation)
Defined as $G_m = (S,\,C,\,U,\,V,\,E)$:
- $S$: finite set of scenarios (decision prompts)
- $C$: set of moral contexts (e.g., “when advising someone in distress”)
- $U$: participant pool
- $V$: deduplicated canonical value cards
- $E \subseteq V \times V \times C$: directed edges $(v_a \to v_b, c)$, indicating that under context $c$, $v_b$ is regarded (by majority “wisdom vote”) as “wiser” than $v_a$ [2404.10636]

### Moral Event Graph (Event Extraction)
Formally, $G = (V, E, A)$:
- $V = V_{ent} \cup V_{evt}$: entity nodes (persons, organizations, GPEs, ‘Other’) and event nodes (moral-action triggers)
- $E \subseteq (V_{ent} \times V_{evt}) \cup (V_{evt} \times V_{ent})$: edges labeled by agent–of (ent→evt) or patient–of (evt→ent)
- $A$: attribute mapping, providing node types, event status (Actual/Intentional/Speculative), MFT moral labels, edge confidence, etc. [2311.09733]

### Moral Preference Graph (LLM Moral Mind)
A probabilistic, network-based structure where nodes are models and edges quantify behavioral moral similarity (based on revealed preference patterns) under moral dilemmas [2412.04476].

## 2. Elicitation and Construction Methodologies

Moral graph elicitation protocols vary by domain (human, text, LLM), but all employ explicit algorithms and structured input types.

### Moral Graph Elicitation (MGE)
A four-phase protocol utilizing LLM-guided interviews [2404.10636]:
1. **Scenario Selection:** Choose morally salient user prompts $S$.
2. **LLM-guided Value Extraction:** Chatbot interviews $u \in U$ using scenario $s \in S$ to surface “constitutive attentional policies” (CAPs), consolidated into editable value cards.
3. **Deduplication:** Canonicalize new value cards via embedding-based similarity and human/LLM validation.
4. **Wisdom-Voting and Edges:** For each context $c \in C$, participants judge LLM-generated transition stories (agent goes from $v_i$ to $v_j$) as wisdom “upgrades.” Affirmative votes instantiate edges $(v_i \to v_j, c)$.
5. **Ranking:** PageRank on context-labeled subgraphs produces a hierarchy of values for each $c$.

**Pseudocode Example** (abridged from [2404.10636]):
```python
function MGE(S, U):
  V, E ← ∅
  for u in U:
    s ← u.select_scenario(S)
    raw_cards ← LLM_interview(u, s)
    for v′ in raw_cards:
      v ← deduplicate(v′, V)
      V.add(v)
  C ← extract_contexts(S)
  for c in C:
    for (v_i, v_j) relevant to c:
      story ← LLM_generate_transition(v_i, v_j, c)
      for u in U:
        vote ← u.judge_wisdom(story)
        if vote == "Yes": E.add((v_i → v_j, c))
  return MoralGraph(S, C, U, V, E)
```

### Moral Event Graph Extraction (MOKA)
End-to-end pipeline [2311.09733]:
1. Preprocess (segment, NER, candidate entity extraction)
2. Event detection (generative or sequence tagging)
3. Argument labeling (extract agent/patient/morality/status)
4. Moral Knowledge Augmentation (lexicon and scenario retrieval)
5. Graph construction (populate $V_{ent}$, $V_{evt}$, $E$, assign $A$)

### Moral Preference Graph Elicitation (LLMs)
“Priced Survey Methodology” [2412.04476]:
1. Structured moral dilemmas over 5 axes, with randomized budget constraints.
2. Each LLM makes 161 discrete choices (1 unconstrained, 160 constrained).
3. Rationality test via $\epsilon$-relaxed GARP and Afriat’s CCEI; pass/fail indicates rational moral agenthood.
4. Fitting single-peaked quadratic utility (per dimension), revealing latent weights and ideal points.
5. Permutation-based MILP clustering yields networks of behavioral moral similarity.

## 3. Six Criteria for an Alignment Target

The MGE framework specifies six desiderata for any moral graph used as an alignment target [2404.10636]:

| Property        | Type                    | Description                                                |
|-----------------|------------------------|------------------------------------------------------------|
| Fine-grained    | Steering-behavior       | Handles specific, granular contexts, not just abstractions  |
| Generalizable   | Steering-behavior       | Elicited values transfer to novel, related situations       |
| Scalable        | Steering-behavior       | More participants improve coverage, not dilute wisdom       |
| Robust          | Political/workability   | Resistant to manipulation or ideological capture            |
| Legitimate      | Political/workability   | Stakeholders recognize and endorse the process and outputs  |
| Auditable       | Political/workability   | Human-readable, explicit, and traceable artifacts           |

These criteria differentiate moral graphs from benchmark-based or distributional targets by making transparency, context-sensitivity, and robustness explicit design objectives.

## 4. Aggregation Algorithms and Graph Analysis

Aggregation over elicited moral data is typically performed using graph-theoretic or network algorithms.

### PageRank-based Aggregation
Given a context-specific subgraph $G_c = (V, E_c)$, PageRank scoring yields
\[
\mathrm{PR}_c(v) = \alpha \sum_{(u\to v)\in E_c} \frac{\mathrm{PR}_c(u)}{\mathrm{outdeg}_c(u)} + \frac{1-\alpha}{|V|}
\]
with $\alpha\approx0.85$ [2404.10636]. High-scoring values (roots of “wisdom” flows) represent the model’s optimal targets for context $c$.

### Non-Parametric Similarity Networks (LLMs)
Moral preference graphs among LLMs are constructed by repeated MILP-based partitioning and permutation testing:
- $G[m,w]$ quantifies probability that models $m$ and $w$ assign to the same “type”
- Adjacency matrix $H^\alpha$ defines edges for statistical similarity thresholds
- Standard measures (betweenness/eigenvector centrality, connected components) reveal core, bridging, and cluster structure [2412.04476]

**Empirical findings**: loose clustering thresholds produce a single “shared core,” while strict thresholds fragment the space into discrete “moral camps” or stances.

## 5. Empirical Applications and Quantitative Outcomes

Moral graph elicitation has been empirically validated in both human and machine contexts.

### Human Elicitation (MGE Study)
- 500 participants (stratified U.S. sample), ~35k GPT-4 tokens per participant, median survey duration 15 minutes [2404.10636]
- Abortion, parenting, and Jan 6 “weapons” scenarios used as cases
- 85 deduplicated value cards, ~100 edges, few cycles
- 89.1% of participants felt well-represented; 89% thought the graph was fair
- Robustness: wisdom edges and value rankings not predictable from initial ideological stance
- Emergent expertise: values articulated by experienced stakeholders (“Informed Autonomy” on abortion) ascend to top ranks via transitive wisdom edges and PageRank

### Moral Event Graphs (MOKA System)
- 5,494 event annotations on 474 news articles, each event mapped to agent, patient, trigger, status, and 10 MFT labels [2311.09733]
- MOKA introduces lexicon and scenario modules, achieving superior event detection and moral role labeling

### LLM Moral Graphs
- 40 leading LLMs, each evaluated over 160+ structured dilemmas [2412.04476]
- At least one model per provider passes probabilistic rationality; single-peaked utilities estimated
- Moral similarity network reveals central “core” group plus clusters of rigid or flexible ethical profiles

## 6. Critical Discussion, Strengths, and Limitations

### Strengths of Moral Graph Elicitation
- **Contextuality & interpretability:** Each edge is labeled by scenario/context; value cards consist of explicit CAPs rather than vague slogans.
- **Scalability:** Transitive wisdom-flow (PageRank, vote aggregation) surfaces nuanced, expert-informed values without the need for predefined expert labels [2404.10636].
- **Empirical robustness:** Outputs are resistant to adversarial gaming and maintain participant legitimacy—confirmed quantitatively [2404.10636].
- **Modularity:** Graph-based outputs can be incorporated directly in downstream alignment (policy/behavioral) pipelines or diagnostic benchmarks [2311.09733].

### Limitations
- **Resource intensity:** Human protocols (~35k LLM tokens per user) and requirement of multi-phase interviews increase computational and participant burden [2404.10636].
- **Generalizability scope:** Studies to date focus on U.S. samples and English-language prompts.
- **Edge-case cycles:** Zero-sum power conflicts or non-transitive preferences can induce graph cycles, sometimes omitted in output (with modeling discretion required) [2404.10636].
- **LLM involvement bias:** LLM-generated transition stories may introduce subtle framing effects in “wisdom-vote” tasks.

### Extensions
- Global implementations with multilingual samples and automated context-detection in downstream applications.
- Hybrid (human+LLM) storytelling and iterative, real-time updating of the moral graph as deployed models collect real-world feedback [2404.10636].

## 7. Connections to Event Extraction, Value Alignment, and AI Ethics

Moral graph elicitation unifies threads across computational ethics:
- As an **alignment primitive**, it directly addresses the “elicitation” and “reconciliation” phases of the AI value-alignment pipeline [2404.10636].
- In **natural language processing**, event-level graph representations (MEGs) operationalize moral judgments in text, leveraging both lexicon- and scenario-based memory [2311.09733].
- As a **diagnostic for model behavior**, preference graphs for LLMs enable empirical measurement of coherence, centrality, and diversity in moral reasoning [2412.04476].

Theoretical underpinnings trace to “wisdom upgrades” (Taylor 1989; Chang 2004), revealed preference theory, and Moral Foundations Theory (MFT), providing both interpretability and mathematical tractability for scaling alignment across diverse populations and model classes.

Source: https://www.emergentmind.com/topics/moral-graph-elicitation