Moral Graph Elicitation
- Moral graph elicitation is a structured approach that uses graph representations to model moral values, agents, and contexts.
- It employs systematic protocols such as LLM-guided interviews, wisdom voting, and PageRank aggregation to rank moral preferences.
- The methodologies support value alignment in AI and natural language processing with robust, scalable, and transparent outputs.
Moral graph elicitation refers to a set of rigorous, graph-structured methodologies for extracting, representing, and reconciling moral values—whether from human populations, news text, or AI systems—using explicit graph schemas and computational workflows. These approaches operationalize values and their relationships as nodes and edges, facilitating subsequent alignment, interpretability, or utility analysis. Moral graph elicitation underpins both the design of scalable value alignment protocols for AI and the systematic study of moral event structure in natural language.
1. Formal Definitions of Moral Graphs
A moral graph is a structured, directed representation encoding values, agents, events, and/or comparative value preferences. Distinct formalizations are found in the literature.
Moral Graph (Human Value Elicitation)
Defined as :
- : finite set of scenarios (decision prompts)
- : set of moral contexts (e.g., “when advising someone in distress”)
- : participant pool
- : deduplicated canonical value cards
- : directed edges , indicating that under context , is regarded (by majority “wisdom vote”) as “wiser” than (Klingefjord et al., 2024)
Moral Event Graph (Event Extraction)
Formally, 0:
- 1: entity nodes (persons, organizations, GPEs, ‘Other’) and event nodes (moral-action triggers)
- 2: edges labeled by agent–of (ent→evt) or patient–of (evt→ent)
- 3: attribute mapping, providing node types, event status (Actual/Intentional/Speculative), MFT moral labels, edge confidence, etc. (Zhang et al., 2023)
Moral Preference Graph (LLM Moral Mind)
A probabilistic, network-based structure where nodes are models and edges quantify behavioral moral similarity (based on revealed preference patterns) under moral dilemmas (Seror, 2024).
2. Elicitation and Construction Methodologies
Moral graph elicitation protocols vary by domain (human, text, LLM), but all employ explicit algorithms and structured input types.
Moral Graph Elicitation (MGE)
A four-phase protocol utilizing LLM-guided interviews (Klingefjord et al., 2024):
- Scenario Selection: Choose morally salient user prompts 4.
- LLM-guided Value Extraction: Chatbot interviews 5 using scenario 6 to surface “constitutive attentional policies” (CAPs), consolidated into editable value cards.
- Deduplication: Canonicalize new value cards via embedding-based similarity and human/LLM validation.
- Wisdom-Voting and Edges: For each context 7, participants judge LLM-generated transition stories (agent goes from 8 to 9) as wisdom “upgrades.” Affirmative votes instantiate edges 0.
- Ranking: PageRank on context-labeled subgraphs produces a hierarchy of values for each 1.
Pseudocode Example (abridged from (Klingefjord et al., 2024)): 5
Moral Event Graph Extraction (MOKA)
End-to-end pipeline (Zhang et al., 2023):
- Preprocess (segment, NER, candidate entity extraction)
- Event detection (generative or sequence tagging)
- Argument labeling (extract agent/patient/morality/status)
- Moral Knowledge Augmentation (lexicon and scenario retrieval)
- Graph construction (populate 2, 3, 4, assign 5)
Moral Preference Graph Elicitation (LLMs)
“Priced Survey Methodology” (Seror, 2024):
- Structured moral dilemmas over 5 axes, with randomized budget constraints.
- Each LLM makes 161 discrete choices (1 unconstrained, 160 constrained).
- Rationality test via 6-relaxed GARP and Afriat’s CCEI; pass/fail indicates rational moral agenthood.
- Fitting single-peaked quadratic utility (per dimension), revealing latent weights and ideal points.
- Permutation-based MILP clustering yields networks of behavioral moral similarity.
3. Six Criteria for an Alignment Target
The MGE framework specifies six desiderata for any moral graph used as an alignment target (Klingefjord et al., 2024):
| Property | Type | Description |
|---|---|---|
| Fine-grained | Steering-behavior | Handles specific, granular contexts, not just abstractions |
| Generalizable | Steering-behavior | Elicited values transfer to novel, related situations |
| Scalable | Steering-behavior | More participants improve coverage, not dilute wisdom |
| Robust | Political/workability | Resistant to manipulation or ideological capture |
| Legitimate | Political/workability | Stakeholders recognize and endorse the process and outputs |
| Auditable | Political/workability | Human-readable, explicit, and traceable artifacts |
These criteria differentiate moral graphs from benchmark-based or distributional targets by making transparency, context-sensitivity, and robustness explicit design objectives.
4. Aggregation Algorithms and Graph Analysis
Aggregation over elicited moral data is typically performed using graph-theoretic or network algorithms.
PageRank-based Aggregation
Given a context-specific subgraph 7, PageRank scoring yields
8
with 9 (Klingefjord et al., 2024). High-scoring values (roots of “wisdom” flows) represent the model’s optimal targets for context 0.
Non-Parametric Similarity Networks (LLMs)
Moral preference graphs among LLMs are constructed by repeated MILP-based partitioning and permutation testing:
- 1 quantifies probability that models 2 and 3 assign to the same “type”
- Adjacency matrix 4 defines edges for statistical similarity thresholds
- Standard measures (betweenness/eigenvector centrality, connected components) reveal core, bridging, and cluster structure (Seror, 2024)
Empirical findings: loose clustering thresholds produce a single “shared core,” while strict thresholds fragment the space into discrete “moral camps” or stances.
5. Empirical Applications and Quantitative Outcomes
Moral graph elicitation has been empirically validated in both human and machine contexts.
Human Elicitation (MGE Study)
- 500 participants (stratified U.S. sample), ~35k GPT-4 tokens per participant, median survey duration 15 minutes (Klingefjord et al., 2024)
- Abortion, parenting, and Jan 6 “weapons” scenarios used as cases
- 85 deduplicated value cards, ~100 edges, few cycles
- 89.1% of participants felt well-represented; 89% thought the graph was fair
- Robustness: wisdom edges and value rankings not predictable from initial ideological stance
- Emergent expertise: values articulated by experienced stakeholders (“Informed Autonomy” on abortion) ascend to top ranks via transitive wisdom edges and PageRank
Moral Event Graphs (MOKA System)
- 5,494 event annotations on 474 news articles, each event mapped to agent, patient, trigger, status, and 10 MFT labels (Zhang et al., 2023)
- MOKA introduces lexicon and scenario modules, achieving superior event detection and moral role labeling
LLM Moral Graphs
- 40 leading LLMs, each evaluated over 160+ structured dilemmas (Seror, 2024)
- At least one model per provider passes probabilistic rationality; single-peaked utilities estimated
- Moral similarity network reveals central “core” group plus clusters of rigid or flexible ethical profiles
6. Critical Discussion, Strengths, and Limitations
Strengths of Moral Graph Elicitation
- Contextuality & interpretability: Each edge is labeled by scenario/context; value cards consist of explicit CAPs rather than vague slogans.
- Scalability: Transitive wisdom-flow (PageRank, vote aggregation) surfaces nuanced, expert-informed values without the need for predefined expert labels (Klingefjord et al., 2024).
- Empirical robustness: Outputs are resistant to adversarial gaming and maintain participant legitimacy—confirmed quantitatively (Klingefjord et al., 2024).
- Modularity: Graph-based outputs can be incorporated directly in downstream alignment (policy/behavioral) pipelines or diagnostic benchmarks (Zhang et al., 2023).
Limitations
- Resource intensity: Human protocols (~35k LLM tokens per user) and requirement of multi-phase interviews increase computational and participant burden (Klingefjord et al., 2024).
- Generalizability scope: Studies to date focus on U.S. samples and English-language prompts.
- Edge-case cycles: Zero-sum power conflicts or non-transitive preferences can induce graph cycles, sometimes omitted in output (with modeling discretion required) (Klingefjord et al., 2024).
- LLM involvement bias: LLM-generated transition stories may introduce subtle framing effects in “wisdom-vote” tasks.
Extensions
- Global implementations with multilingual samples and automated context-detection in downstream applications.
- Hybrid (human+LLM) storytelling and iterative, real-time updating of the moral graph as deployed models collect real-world feedback (Klingefjord et al., 2024).
7. Connections to Event Extraction, Value Alignment, and AI Ethics
Moral graph elicitation unifies threads across computational ethics:
- As an alignment primitive, it directly addresses the “elicitation” and “reconciliation” phases of the AI value-alignment pipeline (Klingefjord et al., 2024).
- In natural language processing, event-level graph representations (MEGs) operationalize moral judgments in text, leveraging both lexicon- and scenario-based memory (Zhang et al., 2023).
- As a diagnostic for model behavior, preference graphs for LLMs enable empirical measurement of coherence, centrality, and diversity in moral reasoning (Seror, 2024).
Theoretical underpinnings trace to “wisdom upgrades” (Taylor 1989; Chang 2004), revealed preference theory, and Moral Foundations Theory (MFT), providing both interpretability and mathematical tractability for scaling alignment across diverse populations and model classes.