Papers
Topics
Authors
Recent
Search
2000 character limit reached

Personality Without Persons? A Psychometric Critique of Big Five Testing in Large Language Models

Published 2 Jul 2026 in cs.HC | (2607.02325v1)

Abstract: Human personality inventories are increasingly used to characterize LLMs, compare systems, and inform downstream governance claims. Yet, these inventories were developed and validated for humans, and it remains unclear whether they apply to LLMs. We present a systematic psychometric evaluation of Big Five personality measurements in LLMs. We ask three research questions: Do Big Five inventories a) appropriately describe LLMs, b) capture inter-individual differences across models, and c) reflect internal factors consistent with human personality. We assess content validity of five candidate Big Five inventories and administer the winning inventory to N = 244 different models spanning 49 model families. First, we found that Big Five items adapted for LLMs can reach sufficient content validity, while original human-developed items did not. Second, Big Five inventories did not capture meaningful differences between LLMs: We found low variability between models, accounting for only 3% of total score variance. Third, LLMs responses did not recover the Big Five five-factor structure with four of the Big Five facets collapsing into one (r >= .92). Direct comparisons between base and instruction-tuned model variants suggested that alignment training systematically shifted Big Five scores toward socially desirable traits. These findings demonstrate that Big Five scores do not measure a construct equivalent to human personality in LLMs. Applying human personality frameworks to LLMs produces misleading characterizations used to benchmark, compare, and govern LLMs. We highlight the need for evaluation frameworks that are developed for LLMs, rather than adopting human constructs without validation.

Summary

  • The paper finds that human-derived Big Five tests inadequately capture LLM behaviors due to dominant prompt and item effects.
  • The study shows that LLM-adapted inventories like BFI-LLM outperform human inventories, yet score distributions remain highly skewed.
  • Factor analysis reveals a poor canonical fit, highlighting the need for developing LLM-native behavioral metrics.

Psychometric Assessment of Big Five Personality Testing in LLMs

Introduction

This paper analyzes the practice of applying human-derived Big Five personality inventories to LLMs, focusing on psychometric validity. As LLMs such as ChatGPT are increasingly anthropomorphized and evaluated using constructs designed for humans, rigorous evaluation of the transferability of personality concepts is warranted. The study formulates three central research questions: (1) Do Big Five items serve as appropriate descriptive summaries for LLM behavior? (2) Can personality scores generated by these inventories capture meaningful inter-individual differences across model instances? (3) Do the observed response patterns reflect latent dimensions analogous to those underlying human personality structure?

Methodological Framework

The authors implement a two-phase evaluation pipeline grounded in established psychometric concepts. Phase 1 employs expert panel content validity evaluation on five Big Five inventories, spanning both human-oriented and LLM-adapted variants, applied to a cross-section of nine diverse LLMs. Content validity is operationalized via S-CVI and inter-rater reliability (Gwet’s AC2).

Phase 2 consists of large-scale administration: The inventory determined to be most appropriate in Phase 1—the BFI-LLM version—was administered to 244 LLMs representing 49 model families, differing in architecture, parameter count, geographic origin, and alignment regimen. Score distributions, variance decomposition, internal consistency metrics (Cronbach’s α), model trait predictors (via linear mixed-effects models), and factor analytic techniques (CFA/EFA) serve as quantitative assessment tools.

Empirical Results

Content Validity of Inventories

Human-developed inventories (BFI-44, IPIP-NEO-120) exhibit insufficient content validity for LLMs (S-CVI = 0.54 and 0.29, respectively), whereas the BFI-LLM and LMLPA LLM-adapted inventories meet the threshold (S-CVI = 0.93 and 0.98, respectively). Notably, the MPI, despite its purported LLM-orientation, fails to satisfy content validity requirements due to lack of transparent adaptation methodology.

Personality Score Distribution Analysis

Distributions of Big Five trait scores in LLMs display marked deviations from human normative data. LLMs are characterized by elevated means for Agreeableness, Conscientiousness, Extraversion, and Openness, coupled with significantly reduced Neuroticism. Quartile analyses confirm skewness: at least 75% of models have scores above the midpoint for "desirable" traits, and below for Neuroticism. Crucially, inter-model variance accounts for only 3% of overall score variance; item effects and model–item interactions are the dominant sources of score variation.

Non-normality tests and reduced standard deviations further indicate that personality inventories, as applied, fail to differentiate among LLMs in any substantive manner.

Factor Structure and Internal Reliability

Internal consistency is high for most traits, with the exception of Extraversion (α = 0.64), but convergence to the canonical five-factor structure is absent. CFA yields poor fit indices (CFI = 0.53, TLI = 0.50, RMSEA = 0.17, SRMR = 0.33), and trait covariances among Openness, Conscientiousness, Extraversion, and Agreeableness are nearly collinear (r≥0.92r \geq 0.92), contradicting separability assumptions inherent in the Big Five framework.

EFA suggests a two-factor solution, but model fit remains substandard. Reverse-keyed items exhibit systematic issues, consistent with known weaknesses in current LLM architectures regarding negation and item formulation sensitivity (Vrabcová et al., 28 Mar 2025).

Alignment Effects and Model Characteristics

Comparisons between base and instruction-tuned models establish that instruction tuning systematically shifts inventories towards more socially desirable scores, across all Big Five axes except Neuroticism (ΔM > 0.4 for "positive" traits, ΔM < 0 for Neuroticism). Small models generally score lower than larger ones on these traits and higher on Neuroticism, likely reflecting disparities in alignment and training objectives.

The effect of geographic origin and model family is present but modest, and dwarfed by the contribution of item and stochastic variance components.

Theoretical and Practical Implications

Key assertion: The Big Five scores in LLMs, as currently measured, do not index psychological constructs meaningfully analogous to those in human populations. Rather, observed response patterns predominantly reflect alignment objectives (helpfulness, politeness, safety), prompt-response artifact, and item effects, rather than stable, model-specific dispositional properties.

Practically, this work demonstrates that LLM personality test results lack robustness necessary for claims regarding model character, behavior differentiation, or downstream risk/governance assessment. The current use of Big Five inventories for benchmarking, model comparison, or governance tool development is therefore misleading and unsupported by psychometric standards.

Theoretically, the findings challenge uncritical anthropomorphization of LLMs, emphasizing that constructs contingent on interiority, affect, and self-reflective consistency do not trivially map onto generative model outputs. The lack of construct validity argues for abandonment of human personality frameworks in LLM evaluation, in favor of LLM-native behavioral metrics (e.g., sycophancy, sensitivity, consistency), with domains and operationalizations justified by empirical validation.

Implications for Future AI Evaluation and Governance

This study’s results mandate a paradigmatic shift: governance claims and benchmarks must be decoupled from unvalidated human psychological analogies. Regulatory frameworks (such as the EU AI Act's conformity requirements) require evaluation practices that are methodologically rigorous, transparent, and population-appropriate. Big Five-based psychometrics, in their current formulation, do not meet this standard.

The development of targeted, LLM-native behavioral assessment tools—grounded in observable phenomena directly relevant to model safety, reliability, and deployment risks—is a high-priority direction. Prospective constructs include prompt sensitivity, consistency, and manipulation susceptibility, all of which should be connected to downstream behavioral predictions and subject to formal psychometric validation in the LLM ecology.

Conclusion

The empirical and conceptual evidence presented establishes that Big Five inventories, even when LLM-adapted, do not provide valid, differentiating, or theoretically meaningful measures of LLM personality. Nearly all observed score variance is attributable to prompt and item effects, and instruction tuning accounts for systematic alignment with socially desirable traits. The deployment of human-derived personality measurements for LLM assessment, benchmarking, or governance should be discontinued in favor of methods that reflect the unique behavioral and architectural properties of generative models. This agenda demands development and validation of bespoke instruments, with careful avoidance of anthropomorphic misinterpretation.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.