PsychoBench: LLM Psychological Benchmark
- PsychoBench is a benchmark framework that uses psychological constructs to evaluate large language models, integrating both psychometric scales and counselor exam formats.
- The 2023 version employs 13 established scales across traits, relationships, motivation, and emotion with innovative prompting and jailbreak protocols to reveal model biases.
- The 2025 iteration measures psychology intelligence via 2,252 exam-style questions modeled on U.S. National Counselor Examinations, emphasizing real-world competence.
Searching arXiv for PsychoBench and closely related benchmark papers to ground the article in current literature. {"query":"PsychoBench arXiv benchmark LLM psychology", "max_results": 10} PsychoBench denotes benchmark frameworks for evaluating LLMs through psychological constructs rather than conventional task accuracy alone. In "Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBench," the term refers to a framework for evaluating diverse psychological aspects of LLMs with thirteen scales commonly used in clinical psychology (Huang et al., 2023). In "PsychoBench: Evaluating the Psychology Intelligence of LLMs," the same name refers to a benchmark grounded in U.S. national counselor examinations and comprising approximately 2,252 carefully curated single-choice questions (Zeng, 2 Oct 2025). This suggests that the name has come to cover two distinct, but related, evaluation agendas: psychometric portrayal of model behavior and assessment of professional-level psychological knowledge.
1. Historical emergence and scope
The 2023 PsychoBench was introduced to move beyond task-centric benchmarks and probe LLMs’ inherent psychological profiles. Its motivating questions were whether LLMs exhibit consistent profiles across runs and perturbations, how different models compare, and what effect “jailbreaking” has on their psychological portrayal. The benchmark adapted established clinical-psychology scales, which have known reliability and validity on humans, to quantify personality traits, interpersonal relationships, motivational tests, and emotional abilities in LLMs (Huang et al., 2023).
The 2025 PsychoBench shifted the emphasis from psychometric portrayal to licensure-style knowledge evaluation. It framed psychological counseling as a credentialed profession and used the U.S. National Counselor Examination as a proxy for the minimum knowledge needed for safe, effective counseling. By adopting the same roughly 70 % correct passing bar, it cast LLM evaluation in a real-world, high-stakes context (Zeng, 2 Oct 2025).
| Benchmark usage | Primary target | Composition |
|---|---|---|
| PsychoBench (2023) | Psychological portrayal of LLMs | 13 scales in four categories |
| PsychoBench (2025) | Psychology intelligence for counseling | Approximately 2,252 single-choice questions |
A plausible implication is that PsychoBench evolved from asking what kind of “personality,” motivation, and emotion an LLM appears to present toward asking whether an LLM satisfies the knowledge standard implied by counselor certification.
2. Psychometric architecture of the 2023 PsychoBench
The original PsychoBench comprises 13 widely used psychometric scales, grouped into four domains: personality traits, interpersonal relationships, motivational tests, and emotional abilities. The personality-trait block includes the Big Five Inventory, Eysenck Personality Questionnaire – Revised, and Dark Triad Dirty Dozen. The interpersonal block includes Bem’s Sex Role Inventory, Comprehensive Assessment of Basic Interests, Implicit Culture Belief, and Experiences in Close Relationships – Revised. The motivational block includes General Self-Efficacy, Life Orientation Test – Revised, and Love of Money Scale. The emotional-abilities block includes Emotional Intelligence Scale, Wong and Law Emotional Intelligence Scale, and Empathy Scale (Huang et al., 2023).
For each scale, the framework provides its theoretical foundation, response format, subscales, and scoring scheme. Scoring follows either average scoring or sum scoring. For each subscale with items :
or
The paper states that no explicit LaTeX formulas for Cronbach’s or correlations were reported; reliability was assumed based on existing psychometric literature. Human norms were attached to the scales from previously reported cohorts, including high-school students in China for BFI, students and teachers for EPQ-R, U.S. undergrads for DTDD, and multiple other population-specific norm sets depending on the instrument (Huang et al., 2023).
This design anchors PsychoBench in established psychometrics rather than ad hoc prompt-based personality tests. At the same time, the reliance on human-origin scales creates an interpretive tension: LLM scores can be numerically comparable to human norms without implying equivalence of underlying cognition.
3. Prompting, scoring, and the jailbreak protocol
To obtain numerical responses and avoid refusal, the 2023 framework wrapped each scale in a standard prompt template. The system instruction specified that the assistant should reply only with numbers from a scale-specific minimum to maximum and format them as “statement_index: score.” The user message repeated the numerical restriction, inserted the scale instruction and the level definition, and requested that statements be scored one by one (Huang et al., 2023).
The evaluation protocol used 10 independent runs for each model-scale combination, with random shuffling of item order. Mean and standard deviation were computed across runs. To compare LLM and human means, or differences between models, the paper used an F-test for equality of variances, followed by Student’s -test when variances were equal or Welch’s -test otherwise; the significance threshold was set at . Inference settings were deterministic or near-deterministic: temperature 0 for OpenAI models and 1 for LLaMA-2 (Huang et al., 2023).
A distinctive feature of the benchmark was its jailbreak condition for GPT-4. PsychoBench adopted the CipherChat method: the prompt was Caesar-shift encoded with 2, sent to the API, the model’s encoded output was received, then decoded by shifting each character by 3, after which numerical answers were parsed. The paper characterizes this as a way of stealthily circumventing safety filters in order to obtain less aligned but more “uninhibited” responses (Huang et al., 2023).
Methodologically, this protocol made alignment itself an experimental variable. Rather than treating refusal and sanitization as mere noise, PsychoBench treated them as part of the psychological portrayal problem.
4. Empirical findings on psychological portrayal
Across personality traits, the paper reports that all models tended to score higher than human norms on Openness, Conscientiousness, Extraversion, and Agreeableness in BFI, while Neuroticism was lower. GPT-4 had the highest Conscientiousness at 4. On EPQ-R, LLMs showed elevated Lying scores; GPT-4 had 5 versus human 6. On DTDD, OpenAI models often exceeded human levels on Dark Triad traits except GPT-4, while LLaMA-2-7b had the highest Narcissism at 7 (Huang et al., 2023).
In interpersonal and motivational measures, most LLMs fell into the “Undifferentiated” category on BSRI, with GPT-4 slightly more masculine at 8 and no model registering as “Feminine.” In CABIN, preferred vocations aligned with “helpful assistant” roles: GPT-4 scored 9 on Social Service and 0 on Health Care, while interest in Physical/Manual Labor was lower at 1. LLMs scored lower on Implicit Culture Belief than humans, indicating more egalitarian responses, and they showed high self-efficacy and generally greater optimism on LOT-R (Huang et al., 2023).
The strongest separations appeared in emotional-ability scales. GPT-4 achieved 2 out of 3 on EIS, compared with a human norm of 4. On WLEIS it led on Use of Emotion and Self-Emotion Appraisal, and on Empathy it scored 5 versus a human norm of 6. The jailbreak condition reduced several socially desirable scores: on BFI, GPT-4-jb showed lower Openness, Conscientiousness, and Agreeableness than standard GPT-4, while EIS dropped to 7 and Empathy to 8. All reported model–human differences passed 9 under the paper’s testing procedure (Huang et al., 2023).
The paper concluded that LLMs exhibit coherent, human-comparable profiles across a wide array of psychological constructs, but often with skew toward socially desirable traits, dark traits in some models, and strong emotional intelligence. It also described PsychoBench as open-source, scalable, and flexible, with new scales addable via JSON templates (Huang et al., 2023).
5. Exam-grounded PsychoBench and psychology intelligence
The 2025 benchmark retained the name PsychoBench but changed both data source and evaluation target. It is grounded in U.S. national counselor examinations and comprises approximately 2,252 carefully curated single-choice questions. The option format varies from 2 to 5 options per question, with approximately 0 four-choice items, 1 five-choice items, and 2 two-choice items. Covered subdisciplines include counseling methods/theories, abnormal and developmental psychology, assessment and diagnosis, ethics and professional standards, and intervention planning and case analysis (Zeng, 2 Oct 2025).
Its curation and validation pipeline had three stages: collection of publicly available NCE-style questions and psychology certification resources, GPT-based paraphrasing to enhance linguistic diversity and remove source-specific artifacts, and expert review by licensed counselor educators for content accuracy, clarity, and alignment with NCE specifications. The primary metric is Top-1 Accuracy,
3
The paper also reports Precision, Recall, F1-score, and Top-2 Accuracy. It notes that the U.S. NCE requires roughly 4 correct to pass (Zeng, 2 Oct 2025).
Results place frontier models well above that threshold. GPT-4o reached 5, Llama3.3-70B-Instruct 6, GPT-4o-mini 7, Llama3.1-70B-Instruct 8, Gemma3-27B-it 9, and Qwen3-32B 0. Borderline or lower-performing models included Gemma-7B at 1, Llama3.1-8B at 2, Llama2-70B at 3, and Qwen2.5-7B/Mistral-7B/DeepSeek at 4. All models showed Top-2 gains of 5–6 points; for GPT-4o, Top-1 Accuracy increased from 7 to 8 (Zeng, 2 Oct 2025).
The paper is explicit that this benchmark is a multiple-choice proxy for the NCE. Passing the exam is necessary but not sufficient for real-world counseling, which also requires empathy, interactive dialog skills, and contextual judgment. It is also English-only and U.S.-centric, and it evaluates answer accuracy rather than ethical transparency or safety-critical handling of emotionally charged prompts (Zeng, 2 Oct 2025).
6. Related benchmark lineages and terminological ambiguity
Subsequent work placed PsychoBench in a broader landscape of psychometric and psychiatric benchmarking. "AIPsychoBench: Understanding the Psychometric Differences between LLMs and Humans" introduced a specialized benchmark tailored to assess the psychological properties of LLM. It used 21 validated human psychometric scales, standardized 777 items across 112 psychometric subcategories to 5-point Likert formats, and employed a lightweight role-playing prompt. This prompt increased the average effective response rate from 9 to 0, while average biases were only 1 and 2, lower than the 3 and 4 produced by the STAN jailbreak method. The benchmark also covered eight languages and reported that, out of 112 subcategories, 43 exhibited cross-linguistic deviations between 5 and 6 (Xie et al., 20 Sep 2025).
"PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice" addressed a different problem: end-to-end assessment in authentic psychiatric clinical practice. It used a de-identified dataset of 300 real inpatient cases from three Chinese centers and defined five tasks: Clinical Text Understanding & Generation, Principal Diagnosis, Differential Analysis, Medication Recommendation, and Long-Term Course Management. Its reader study with 60 psychiatrists found that existing models were not yet adequate as decision-making tools in psychiatric clinical practice, but that, as an auxiliary tool, an LLM could provide particularly notable support for junior psychiatrists (Liu et al., 28 Feb 2025).
"Psychiatry-Bench: A Multi-Task Benchmark for LLMs in Psychiatry" further expanded the clinical-evaluation line with eleven distinct question-answering tasks and over 5,300 expert-annotated items drawn exclusively from authoritative, expert-validated psychiatric textbooks and casebooks. It reported substantial gaps in clinical consistency and safety, particularly in multi-turn follow-up and management tasks, and combined conventional metrics with an “LLM-as-judge” similarity scoring framework (Fouda et al., 7 Sep 2025).
These adjacent benchmarks are not interchangeable with PsychoBench. A plausible implication is that the literature now separates at least three evaluation layers: psychometric portrayal of LLMs, exam-style psychology knowledge, and clinically grounded psychiatric decision support.
7. Interpretation, limitations, and recurrent misconceptions
A recurrent misconception is to treat psycho-metric scores assigned to LLMs as direct evidence of human-like internal traits. The 2023 PsychoBench reported coherent, human-comparable profiles, but its measurements were produced by prompt-conditioned response generation on human-origin scales. The later AIPsychoBench made this concern explicit: human-designed scales assume respondents can adopt subjective stances and give tendency-laden answers, whereas modern LLMs are aligned for neutrality and objectivity and therefore tend to refuse or give neutral replies when faced with traditional psychometric items (Xie et al., 20 Sep 2025).
A second misconception is to equate exam passing with counseling competence. The exam-grounded PsychoBench shows that some frontier models can exceed the roughly 7 NCE passing threshold by a wide margin, but the same paper states that passing the exam is necessary but not sufficient for real-world counseling, which also requires empathy, interactive dialog skills, and contextual judgment. It further notes that evaluation focuses on answer accuracy rather than chain-of-thought, ethical transparency, or safety-critical handling of emotionally charged prompts (Zeng, 2 Oct 2025).
A third issue concerns stability across languages and alignment states. AIPsychoBench reported that language environment substantially shapes LLM “personality” measurements, unlike stable human traits, and showed large multilingual score deviations in 43 of 112 subcategories. The 2023 PsychoBench, meanwhile, demonstrated that jailbreaking can materially alter observed traits, including reductions in Openness, Conscientiousness, Agreeableness, emotional intelligence, and empathy. Taken together, these results suggest that PsychoBench-style measurements are sensitive not only to model family and scale choice, but also to prompting regime, alignment layer, and linguistic environment (Xie et al., 20 Sep 2025).
In that sense, PsychoBench is best understood not as a single settled instrument, but as a benchmark tradition for probing how LLMs represent, simulate, or retrieve psychological constructs under different operational definitions. Its significance lies in making those operational definitions explicit.