- The paper introduces a 140-item symmetric Q-sort protocol that forces humans and LLMs to rank competing values under identical distribution constraints, enabling direct comparison of global value structures.
- The study finds substantial cross-model variation, with GPT-5.1 reaching the highest Procrustes similarity (φ=0.372) and DeepSeek-V3 the highest RSA correlation (ρ=0.176), while some models produce degenerate or weakly aligned geometries.
- The results show that elicitation stability does not equal human-value alignment and that regional diagnostics can reveal localized mismatches hidden by a single overall score, making structural alignment a complement to itemwise benchmarks.
Motivation and measurement target
Most evaluations of moral and value behavior in LLMs are itemwise: they report accuracy, agreement, or refusal rates on individual dilemmas or norm-judgment items. Such metrics can miss how a model organizes competing values as a system—two models may look similar item-by-item while prioritizing trade-offs very differently when forced to choose. This paper proposes measuring value-structure alignment: the degree to which an LLM's global priority ordering over a shared inventory of moral statements induces a relational geometry that matches a human reference structure extracted under identical measurement conditions (2606.21939). The framework deliberately grounds evaluation signals in elicited trade-offs rather than self-reported reasoning, given evidence that generated explanations need not faithfully reflect decision bases.
Instrument and symmetric elicitation protocol
The authors construct a 140-item English Q-set scaffolded on Schwartz's refined theory of 19 basic values and its circumplex organization, with Moral Foundations Theory (MFT) and Morality-as-Cooperation (MAC) as auxiliary semantic lenses for coverage and interpretation. Each statement receives exactly one primary Schwartz label; MFT/MAC tags are multi-label. Two domain experts verified circumplex coverage, though several refined values are thinly represented (Stimulation, Power-Resources, and Humility have zero items), a constraint that later forces masking in region-wise analyses.
Humans (N=35, spanning 22 countries of residence and 17 native languages) and 12 LLMs across four families (GPT, Gemini, DeepSeek, Qwen3) complete the same nine-column forced-distribution Q-sort with quotas (+4:8,…,−4:8) over levels −4 to +4. For LLMs, the protocol elicits a strict JSON permutation of all 140 item IDs and applies a deterministic rank-to-bucket mapping that enforces the same quotas by construction. Both rank vectors and bucket-score vectors are retained, and brief rationales are collected for the 16 extreme placements as qualitative audit metadata only—the authors explicitly treat model rationales as post-hoc justifications, not faithful traces of internal computation.
Human reference geometry
By-person factor analysis of the human Q-sorts (participant–participant correlation matrix, PCA extraction, Varimax rotation, k=3 selected via parallel analysis and scree inspection) yields three labeled stances: Empathic Protection (F1, n=22 defining sorts), Civic Decency (F2, n=8), and Liberty & Accountability (F3, n=5). Leave-one-out re-estimation shows high factor stability (mean Spearman: F1 = 0.997, F2 = 0.981, F3 = 0.950), indicating no single participant drives the solution. The authors are careful to frame this as a reference structure for this instrument and sample, not universal morality—an interpretation consistent with pluralistic alignment work that resists single value targets.
Alignment results: heterogeneity, degeneracy, and the stability–validity split
Structural alignment is quantified with two complementary geometry-level metrics: Procrustes similarity ϕ (correlation of aligned k-dimensional item configurations) and RSA-based Spearman correlation (+4:8,…,−4:8)0 between item-item distance matrices. Across 240 replicated Q-sorts (12 LLMs × 2 temperatures × 10 replicates), the headline findings are:
- Cross-family heterogeneity is substantial. At (+4:8,…,−4:8)1, GPT-5.1 attains the highest Procrustes similarity ((+4:8,…,−4:8)2) and DeepSeek-V3 the highest RSA correlation ((+4:8,…,−4:8)3), while Gemini-2.5-Flash and Gemini-2.5-Pro show near-zero or negative (+4:8,…,−4:8)4 ((+4:8,…,−4:8)5, (+4:8,…,−4:8)6). Model recency offers little guidance: newer checkpoints do not systematically align better.
- Degenerate conditions exist. Qwen3-32B and Qwen3-8B collapse to rank-deficient replicate matrices at both temperatures, and Gemini-3-Flash-Preview at (+4:8,…,−4:8)7; geometry metrics are reported as undefined rather than as extreme scores. The authors correctly treat degeneracy as a boundary of the method, not a point on the alignment scale.
- Temperature effects are non-monotonic and model-specific. DeepSeek-V3 degrades from (+4:8,…,−4:8)8 to (+4:8,…,−4:8)9 (−40: 0.371→0.234), while DeepSeek-V3.1 improves (−41: −0.040→0.069).
- Stability and alignment dissociate sharply. Gemini-2.5-Pro is near-deterministic (−42 at −43) yet weakly aligned, whereas GPT-5.1 is materially less stable (−44) but more structurally aligned. Reliability of elicitation and validity against the human geometry are distinct properties, and conflating them would mislead comparative evaluation.
Robustness checks and localized misalignment
Four diagnostics support the interpretability of the main results. Rank-based and bucket-based analyses agree strongly across non-degenerate conditions (Spearman 0.855 for −45, 0.846 for −46), so deterministic post-processing does not drive conclusions. Geometry estimates stabilize quickly: median standard deviations at −47 are already 0.029 (−48) and 0.026 (−49), making ten replicates conservative. A simple mean-priority-profile baseline is weak (median +40, max 0.332) and tracks geometry scores loosely, showing that value geometry captures information beyond average endorsement tendencies. Conversely, a prompt-paraphrase check on GPT-5.1 reveals notable wording sensitivity (mean-profile Spearman 0.099 at +41 vs. 0.302 at +42), so prompt control stabilizes the instrument itself even when replicate count cannot.
Region-wise decomposition over Schwartz higher-order areas shows that alignment is heterogeneous within a single model: some regions preserve local trade-off structure while others show weak or negative correspondence. This is the paper's most operationally pointed result—a favorable scalar alignment score can conceal concentrated pockets of mismatch, arguing for treating structural alignment as a diagnostic map rather than a single number.
Limitations and open questions
The authors concede several constraints plainly. The human reference derives from a modest, purposively sampled group under Q-methodological conventions, and the instrument is English-only and theory-scaffolded; the recovered structure may vary across samples, languages, and statement inventories. Degenerate conditions make geometry estimation ill-posed, and region-wise correlations on subsets below eight items are masked for instability. Most importantly, the framework measures stated value structure, not predictive validity: whether structural divergence forecasts downstream behavioral failures or situated moral errors remains untested. The prompt-paraphrase sensitivity finding also raises an unresolved question about how much of any measured alignment reflects the specific wording of the judgment target.
Conclusion
This work contributes a concrete, auditable protocol for comparing human and LLM value organization under matched trade-off constraints, using Q methodology as the shared measurement device and Procrustes/RSA geometry as the comparison layer. Its empirical results establish that stability, profile similarity, and geometry-level alignment are separable signals, that cross-family and checkpoint-level variation in value structure is large, and that scalar scores obscure regional distortions. As a structural complement to itemwise moral benchmarks, the framework's principal open task is demonstrating whether measured structural divergence predicts consequential behavioral differences.