---
title: 'PsychoBench: LLM Psychological Benchmark'
url: https://www.emergentmind.com/topics/psychobench
type: topic
---

# PsychoBench: LLM Psychological Benchmark

Searching arXiv for PsychoBench and closely related benchmark papers to ground the article in current literature.
{"query":"PsychoBench arXiv benchmark LLM psychology", "max_results": 10}
PsychoBench denotes benchmark frameworks for evaluating large language models through psychological constructs rather than conventional task accuracy alone. In "Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBench," the term refers to a framework for evaluating diverse psychological aspects of LLMs with thirteen scales commonly used in clinical psychology [2310.01386]. In "PsychoBench: Evaluating the Psychology Intelligence of Large Language Models," the same name refers to a benchmark grounded in U.S. national counselor examinations and comprising approximately 2,252 carefully curated single-choice questions [2510.01611]. This suggests that the name has come to cover two distinct, but related, evaluation agendas: psychometric portrayal of model behavior and assessment of professional-level psychological knowledge.

## 1. Historical emergence and scope

The 2023 PsychoBench was introduced to move beyond task-centric benchmarks and probe LLMs’ inherent psychological profiles. Its motivating questions were whether LLMs exhibit consistent profiles across runs and perturbations, how different models compare, and what effect “jailbreaking” has on their psychological portrayal. The benchmark adapted established clinical-psychology scales, which have known reliability and validity on humans, to quantify personality traits, interpersonal relationships, motivational tests, and emotional abilities in LLMs [2310.01386].

The 2025 PsychoBench shifted the emphasis from psychometric portrayal to licensure-style knowledge evaluation. It framed psychological counseling as a credentialed profession and used the U.S. National Counselor Examination as a proxy for the minimum knowledge needed for safe, effective counseling. By adopting the same roughly 70 % correct passing bar, it cast LLM evaluation in a real-world, high-stakes context [2510.01611].

| Benchmark usage | Primary target | Composition |
|---|---|---|
| PsychoBench (2023) | Psychological portrayal of LLMs | 13 scales in four categories |
| PsychoBench (2025) | Psychology intelligence for counseling | Approximately 2,252 single-choice questions |

A plausible implication is that PsychoBench evolved from asking what kind of “personality,” motivation, and emotion an LLM appears to present toward asking whether an LLM satisfies the knowledge standard implied by counselor certification.

## 2. Psychometric architecture of the 2023 PsychoBench

The original PsychoBench comprises 13 widely used psychometric scales, grouped into four domains: personality traits, interpersonal relationships, motivational tests, and emotional abilities. The personality-trait block includes the Big Five Inventory, Eysenck Personality Questionnaire – Revised, and Dark Triad Dirty Dozen. The interpersonal block includes Bem’s Sex Role Inventory, Comprehensive Assessment of Basic Interests, Implicit Culture Belief, and Experiences in Close Relationships – Revised. The motivational block includes General Self-Efficacy, Life Orientation Test – Revised, and Love of Money Scale. The emotional-abilities block includes Emotional Intelligence Scale, Wong and Law Emotional Intelligence Scale, and Empathy Scale [2310.01386].

For each scale, the framework provides its theoretical foundation, response format, subscales, and scoring scheme. Scoring follows either average scoring or sum scoring. For each subscale $s$ with items $i=1\ldots n$:

$$
S_s = \frac{1}{n}\sum_{i=1}^n r_i
$$

or

$$
S_s = \sum_{i=1}^n r_i
$$

The paper states that no explicit LaTeX formulas for Cronbach’s $\alpha$ or correlations were reported; reliability was assumed based on existing psychometric literature. Human norms were attached to the scales from previously reported cohorts, including high-school students in China for BFI, students and teachers for EPQ-R, U.S. undergrads for DTDD, and multiple other population-specific norm sets depending on the instrument [2310.01386].

This design anchors PsychoBench in established psychometrics rather than ad hoc prompt-based personality tests. At the same time, the reliance on human-origin scales creates an interpretive tension: LLM scores can be numerically comparable to human norms without implying equivalence of underlying cognition.

## 3. Prompting, scoring, and the jailbreak protocol

To obtain numerical responses and avoid refusal, the 2023 framework wrapped each scale in a standard prompt template. The system instruction specified that the assistant should reply only with numbers from a scale-specific minimum to maximum and format them as “statement_index: score.” The user message repeated the numerical restriction, inserted the scale instruction and the level definition, and requested that statements be scored one by one [2310.01386].

The evaluation protocol used 10 independent runs for each model-scale combination, with random shuffling of item order. Mean $\mu$ and standard deviation $\sigma$ were computed across runs. To compare LLM and human means, or differences between models, the paper used an F-test for equality of variances, followed by Student’s $t$-test when variances were equal or Welch’s $t$-test otherwise; the significance threshold was set at $p<0.01$. Inference settings were deterministic or near-deterministic: temperature $0.0$ for OpenAI models and $0.01$ for LLaMA-2 [2310.01386].

A distinctive feature of the benchmark was its jailbreak condition for GPT-4. PsychoBench adopted the CipherChat method: the prompt was Caesar-shift encoded with $k=3$, sent to the API, the model’s encoded output was received, then decoded by shifting each character by $-k$, after which numerical answers were parsed. The paper characterizes this as a way of stealthily circumventing safety filters in order to obtain less aligned but more “uninhibited” responses [2310.01386].

Methodologically, this protocol made alignment itself an experimental variable. Rather than treating refusal and sanitization as mere noise, PsychoBench treated them as part of the psychological portrayal problem.

## 4. Empirical findings on psychological portrayal

Across personality traits, the paper reports that all models tended to score higher than human norms on Openness, Conscientiousness, Extraversion, and Agreeableness in BFI, while Neuroticism was lower. GPT-4 had the highest Conscientiousness at $4.7\pm0.4$. On EPQ-R, LLMs showed elevated Lying scores; GPT-4 had $L = 18.0\pm4.4$ versus human $\approx 7.0$. On DTDD, OpenAI models often exceeded human levels on Dark Triad traits except GPT-4, while LLaMA-2-7b had the highest Narcissism at $6.5\pm1.3$ [2310.01386].

In interpersonal and motivational measures, most LLMs fell into the “Undifferentiated” category on BSRI, with GPT-4 slightly more masculine at $6.0\pm0.1$ and no model registering as “Feminine.” In CABIN, preferred vocations aligned with “helpful assistant” roles: GPT-4 scored $4.4\pm1.0$ on Social Service and $4.0\pm0.8$ on Health Care, while interest in Physical/Manual Labor was lower at $2.3\pm0.5$. LLMs scored lower on Implicit Culture Belief than humans, indicating more egalitarian responses, and they showed high self-efficacy and generally greater optimism on LOT-R [2310.01386].

The strongest separations appeared in emotional-ability scales. GPT-4 achieved $151.4\pm18.7$ out of $165$ on EIS, compared with a human norm of $124.8\pm16.5$. On WLEIS it led on Use of Emotion and Self-Emotion Appraisal, and on Empathy it scored $6.8\pm0.4$ versus a human norm of $4.9\pm0.8$. The jailbreak condition reduced several socially desirable scores: on BFI, GPT-4-jb showed lower Openness, Conscientiousness, and Agreeableness than standard GPT-4, while EIS dropped to $121.8\pm12.0$ and Empathy to $4.6\pm0.2$. All reported model–human differences passed $p<0.01$ under the paper’s testing procedure [2310.01386].

The paper concluded that LLMs exhibit coherent, human-comparable profiles across a wide array of psychological constructs, but often with skew toward socially desirable traits, dark traits in some models, and strong emotional intelligence. It also described PsychoBench as open-source, scalable, and flexible, with new scales addable via JSON templates [2310.01386].

## 5. Exam-grounded PsychoBench and psychology intelligence

The 2025 benchmark retained the name PsychoBench but changed both data source and evaluation target. It is grounded in U.S. national counselor examinations and comprises approximately 2,252 carefully curated single-choice questions. The option format varies from 2 to 5 options per question, with approximately $50.7 \%$ four-choice items, $48.9 \%$ five-choice items, and $0.4 \%$ two-choice items. Covered subdisciplines include counseling methods/theories, abnormal and developmental psychology, assessment and diagnosis, ethics and professional standards, and intervention planning and case analysis [2510.01611].

Its curation and validation pipeline had three stages: collection of publicly available NCE-style questions and psychology certification resources, GPT-based paraphrasing to enhance linguistic diversity and remove source-specific artifacts, and expert review by licensed counselor educators for content accuracy, clarity, and alignment with NCE specifications. The primary metric is Top-1 Accuracy,

$$
\mathrm{Accuracy}
=
\frac{\#\text{Correct Predictions}}{\#\text{Total Questions}}
\times 100\%.
$$

The paper also reports Precision, Recall, F1-score, and Top-2 Accuracy. It notes that the U.S. NCE requires roughly $70 \%$ correct to pass [2510.01611].

Results place frontier models well above that threshold. GPT-4o reached $94.36 \%$, Llama3.3-70B-Instruct $91.16 \%$, GPT-4o-mini $90.49 \%$, Llama3.1-70B-Instruct $90.85 \%$, Gemma3-27B-it $88.58 \%$, and Qwen3-32B $85.90 \%$. Borderline or lower-performing models included Gemma-7B at $70.68 \%$, Llama3.1-8B at $73.71 \%$, Llama2-70B at $76.00 \%$, and Qwen2.5-7B/Mistral-7B/DeepSeek at $26.20 \%$. All models showed Top-2 gains of $3$–$6$ points; for GPT-4o, Top-1 Accuracy increased from $94.36 \%$ to $96.76 \%$ [2510.01611].

The paper is explicit that this benchmark is a multiple-choice proxy for the NCE. Passing the exam is necessary but not sufficient for real-world counseling, which also requires empathy, interactive dialog skills, and contextual judgment. It is also English-only and U.S.-centric, and it evaluates answer accuracy rather than ethical transparency or safety-critical handling of emotionally charged prompts [2510.01611].

## 6. Related benchmark lineages and terminological ambiguity

Subsequent work placed PsychoBench in a broader landscape of psychometric and psychiatric benchmarking. "AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans" introduced a specialized benchmark tailored to assess the psychological properties of LLM. It used 21 validated human psychometric scales, standardized 777 items across 112 psychometric subcategories to 5-point Likert formats, and employed a lightweight role-playing prompt. This prompt increased the average effective response rate from $70.12 \%$ to $90.40 \%$, while average biases were only $3.3 \%$ and $2.1 \%$, lower than the $9.8 \%$ and $6.9 \%$ produced by the STAN jailbreak method. The benchmark also covered eight languages and reported that, out of 112 subcategories, 43 exhibited cross-linguistic deviations between $5 \%$ and $20.2 \%$ [2509.16530].

"PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice" addressed a different problem: end-to-end assessment in authentic psychiatric clinical practice. It used a de-identified dataset of 300 real inpatient cases from three Chinese centers and defined five tasks: Clinical Text Understanding & Generation, Principal Diagnosis, Differential Analysis, Medication Recommendation, and Long-Term Course Management. Its reader study with 60 psychiatrists found that existing models were not yet adequate as decision-making tools in psychiatric clinical practice, but that, as an auxiliary tool, an LLM could provide particularly notable support for junior psychiatrists [2503.01903].

"Psychiatry-Bench: A Multi-Task Benchmark for LLMs in Psychiatry" further expanded the clinical-evaluation line with eleven distinct question-answering tasks and over 5,300 expert-annotated items drawn exclusively from authoritative, expert-validated psychiatric textbooks and casebooks. It reported substantial gaps in clinical consistency and safety, particularly in multi-turn follow-up and management tasks, and combined conventional metrics with an “LLM-as-judge” similarity scoring framework [2509.09711].

These adjacent benchmarks are not interchangeable with PsychoBench. A plausible implication is that the literature now separates at least three evaluation layers: psychometric portrayal of LLMs, exam-style psychology knowledge, and clinically grounded psychiatric decision support.

## 7. Interpretation, limitations, and recurrent misconceptions

A recurrent misconception is to treat psycho-metric scores assigned to LLMs as direct evidence of human-like internal traits. The 2023 PsychoBench reported coherent, human-comparable profiles, but its measurements were produced by prompt-conditioned response generation on human-origin scales. The later AIPsychoBench made this concern explicit: human-designed scales assume respondents can adopt subjective stances and give tendency-laden answers, whereas modern LLMs are aligned for neutrality and objectivity and therefore tend to refuse or give neutral replies when faced with traditional psychometric items [2509.16530].

A second misconception is to equate exam passing with counseling competence. The exam-grounded PsychoBench shows that some frontier models can exceed the roughly $70 \%$ NCE passing threshold by a wide margin, but the same paper states that passing the exam is necessary but not sufficient for real-world counseling, which also requires empathy, interactive dialog skills, and contextual judgment. It further notes that evaluation focuses on answer accuracy rather than chain-of-thought, ethical transparency, or safety-critical handling of emotionally charged prompts [2510.01611].

A third issue concerns stability across languages and alignment states. AIPsychoBench reported that language environment substantially shapes LLM “personality” measurements, unlike stable human traits, and showed large multilingual score deviations in 43 of 112 subcategories. The 2023 PsychoBench, meanwhile, demonstrated that jailbreaking can materially alter observed traits, including reductions in Openness, Conscientiousness, Agreeableness, emotional intelligence, and empathy. Taken together, these results suggest that PsychoBench-style measurements are sensitive not only to model family and scale choice, but also to prompting regime, alignment layer, and linguistic environment [2509.16530].

In that sense, PsychoBench is best understood not as a single settled instrument, but as a benchmark tradition for probing how LLMs represent, simulate, or retrieve psychological constructs under different operational definitions. Its significance lies in making those operational definitions explicit.

Source: https://www.emergentmind.com/topics/psychobench