- The paper finds that human-derived Big Five tests inadequately capture LLM behaviors due to dominant prompt and item effects.
- The study shows that LLM-adapted inventories like BFI-LLM outperform human inventories, yet score distributions remain highly skewed.
- Factor analysis reveals a poor canonical fit, highlighting the need for developing LLM-native behavioral metrics.
Psychometric Assessment of Big Five Personality Testing in LLMs
Introduction
This paper analyzes the practice of applying human-derived Big Five personality inventories to LLMs, focusing on psychometric validity. As LLMs such as ChatGPT are increasingly anthropomorphized and evaluated using constructs designed for humans, rigorous evaluation of the transferability of personality concepts is warranted. The study formulates three central research questions: (1) Do Big Five items serve as appropriate descriptive summaries for LLM behavior? (2) Can personality scores generated by these inventories capture meaningful inter-individual differences across model instances? (3) Do the observed response patterns reflect latent dimensions analogous to those underlying human personality structure?
Methodological Framework
The authors implement a two-phase evaluation pipeline grounded in established psychometric concepts. Phase 1 employs expert panel content validity evaluation on five Big Five inventories, spanning both human-oriented and LLM-adapted variants, applied to a cross-section of nine diverse LLMs. Content validity is operationalized via S-CVI and inter-rater reliability (Gwet’s AC2).
Phase 2 consists of large-scale administration: The inventory determined to be most appropriate in Phase 1—the BFI-LLM version—was administered to 244 LLMs representing 49 model families, differing in architecture, parameter count, geographic origin, and alignment regimen. Score distributions, variance decomposition, internal consistency metrics (Cronbach’s α), model trait predictors (via linear mixed-effects models), and factor analytic techniques (CFA/EFA) serve as quantitative assessment tools.
Empirical Results
Content Validity of Inventories
Human-developed inventories (BFI-44, IPIP-NEO-120) exhibit insufficient content validity for LLMs (S-CVI = 0.54 and 0.29, respectively), whereas the BFI-LLM and LMLPA LLM-adapted inventories meet the threshold (S-CVI = 0.93 and 0.98, respectively). Notably, the MPI, despite its purported LLM-orientation, fails to satisfy content validity requirements due to lack of transparent adaptation methodology.
Personality Score Distribution Analysis
Distributions of Big Five trait scores in LLMs display marked deviations from human normative data. LLMs are characterized by elevated means for Agreeableness, Conscientiousness, Extraversion, and Openness, coupled with significantly reduced Neuroticism. Quartile analyses confirm skewness: at least 75% of models have scores above the midpoint for "desirable" traits, and below for Neuroticism. Crucially, inter-model variance accounts for only 3% of overall score variance; item effects and model–item interactions are the dominant sources of score variation.
Non-normality tests and reduced standard deviations further indicate that personality inventories, as applied, fail to differentiate among LLMs in any substantive manner.
Factor Structure and Internal Reliability
Internal consistency is high for most traits, with the exception of Extraversion (α = 0.64), but convergence to the canonical five-factor structure is absent. CFA yields poor fit indices (CFI = 0.53, TLI = 0.50, RMSEA = 0.17, SRMR = 0.33), and trait covariances among Openness, Conscientiousness, Extraversion, and Agreeableness are nearly collinear (r≥0.92), contradicting separability assumptions inherent in the Big Five framework.
EFA suggests a two-factor solution, but model fit remains substandard. Reverse-keyed items exhibit systematic issues, consistent with known weaknesses in current LLM architectures regarding negation and item formulation sensitivity (Vrabcová et al., 28 Mar 2025).
Alignment Effects and Model Characteristics
Comparisons between base and instruction-tuned models establish that instruction tuning systematically shifts inventories towards more socially desirable scores, across all Big Five axes except Neuroticism (ΔM > 0.4 for "positive" traits, ΔM < 0 for Neuroticism). Small models generally score lower than larger ones on these traits and higher on Neuroticism, likely reflecting disparities in alignment and training objectives.
The effect of geographic origin and model family is present but modest, and dwarfed by the contribution of item and stochastic variance components.
Theoretical and Practical Implications
Key assertion: The Big Five scores in LLMs, as currently measured, do not index psychological constructs meaningfully analogous to those in human populations. Rather, observed response patterns predominantly reflect alignment objectives (helpfulness, politeness, safety), prompt-response artifact, and item effects, rather than stable, model-specific dispositional properties.
Practically, this work demonstrates that LLM personality test results lack robustness necessary for claims regarding model character, behavior differentiation, or downstream risk/governance assessment. The current use of Big Five inventories for benchmarking, model comparison, or governance tool development is therefore misleading and unsupported by psychometric standards.
Theoretically, the findings challenge uncritical anthropomorphization of LLMs, emphasizing that constructs contingent on interiority, affect, and self-reflective consistency do not trivially map onto generative model outputs. The lack of construct validity argues for abandonment of human personality frameworks in LLM evaluation, in favor of LLM-native behavioral metrics (e.g., sycophancy, sensitivity, consistency), with domains and operationalizations justified by empirical validation.
Implications for Future AI Evaluation and Governance
This study’s results mandate a paradigmatic shift: governance claims and benchmarks must be decoupled from unvalidated human psychological analogies. Regulatory frameworks (such as the EU AI Act's conformity requirements) require evaluation practices that are methodologically rigorous, transparent, and population-appropriate. Big Five-based psychometrics, in their current formulation, do not meet this standard.
The development of targeted, LLM-native behavioral assessment tools—grounded in observable phenomena directly relevant to model safety, reliability, and deployment risks—is a high-priority direction. Prospective constructs include prompt sensitivity, consistency, and manipulation susceptibility, all of which should be connected to downstream behavioral predictions and subject to formal psychometric validation in the LLM ecology.
Conclusion
The empirical and conceptual evidence presented establishes that Big Five inventories, even when LLM-adapted, do not provide valid, differentiating, or theoretically meaningful measures of LLM personality. Nearly all observed score variance is attributable to prompt and item effects, and instruction tuning accounts for systematic alignment with socially desirable traits. The deployment of human-derived personality measurements for LLM assessment, benchmarking, or governance should be discontinued in favor of methods that reflect the unique behavioral and architectural properties of generative models. This agenda demands development and validation of bespoke instruments, with careful avoidance of anthropomorphic misinterpretation.