---
title: 'PsychoLexEval: Bilingual Psychology Benchmark'
url: https://www.emergentmind.com/topics/psycholexeval
type: topic
---

# PsychoLexEval: Bilingual Psychology Benchmark

PsychoLexEval is a domain-specific benchmark for evaluating psychological knowledge and reasoning in large language models, introduced within the broader PsychoLex resource suite alongside PsychoLexQA and PsychoLexLLaMA. In its original presentation, it is a bilingual Persian–English multiple-choice question and answer benchmark with 3,430 items, designed to provide rigorous, domain-specific evaluation across complex psychological scenarios rather than generic language understanding alone [2408.08848].

## 1. Definition and provenance

PsychoLexEval was proposed to address a gap in psychology-focused evaluation resources for LLMs. The benchmark is described as a novel, rigorously-constructed evaluation dataset intended to test psychological knowledge and reasoning capacity, with explicit bilingual coverage in Persian and English. Its role within the PsychoLex suite is evaluative rather than instructional: PsychoLexQA supplies instructional content, PsychoLexEval supplies benchmarked assessment, and PsychoLexLLaMA represents the specialized model family evaluated against it [2408.08848].

A later line of work places PsychoLexEval at the beginning of the PsychoLexTherapy pipeline, where it is treated as the foundational stage for screening small language models before they are used in reasoning-oriented psychotherapeutic dialogue. In that context, the benchmark is characterized as a curated, domain-specific MCQ dataset in Persian intended to probe whether candidate models possess sufficient psychological knowledge to support downstream therapeutic simulation [2510.03913].

This dual description suggests two closely related perspectives on PsychoLexEval. In the original PsychoLex paper, it functions as a bilingual benchmark for broad evaluation. In PsychoLexTherapy, it is operationalized as a Persian screening instrument for model selection. A plausible implication is that the benchmark serves both as a general research testbed and as a practical gatekeeping mechanism in downstream mental-health-oriented systems.

## 2. Dataset structure and topical coverage

PsychoLexEval is structured as a multiple-choice question and answer dataset. The reported format is MCQA, with exactly four answer options per item and one correct answer. The benchmark contains 3,430 items. The original paper explicitly presents it as bilingual, covering Persian and English [2408.08848].

Its topical scope spans a broad range of psychological disciplines. Reported coverage includes general psychology, developmental psychology, clinical psychology, psychometrics, cognitive testing, industrial/organizational psychology, social psychology, educational psychology, and exceptional children. The benchmark also includes core concepts drawn from introductory psychology, such as biological bases, perception, memory, learning, motivation, emotion, personality, and psychological disorders [2408.08848].

The later PsychoLexTherapy description emphasizes the same broad subfield coverage in Persian, summarizing it as clinical, cognitive, developmental, social, and related areas. That use case frames the benchmark as culturally grounded and language-specific for Persian SLM evaluation [2510.03913].

An example item reported in English is:

> Which strategy is NOT considered a form of problem-focused coping?  
> 1) Defining the problem  
> 2) Seeking emotional support  
> 3) Generating alternative solutions  
> 4) Changing personal goals  
> Correct Answer: 2

This example is representative of the benchmark’s exam-style orientation: the task is discrete, knowledge-centered, and suitable for controlled comparison across prompting conditions and model families [2408.08848].

## 3. Data sources and curation procedures

The original PsychoLex paper reports four main sources for PsychoLexEval: psychology entrance exams from 2014–2024, employment or job qualification tests including psychology-specific ones, reputable online psychology test websites, and GPT-4 generated questions grounded in psychology textbook content. The curation process includes expert review for clarity, relevance, and accuracy; retention only of questions with exactly four options; removal of questions that could pose copyright issues; and rigorous filtering for quality, completeness, and comprehensibility [2408.08848].

The later PsychoLexTherapy description reports a closely aligned construction pipeline for the Persian screening usage: graduate entrance exams in psychology, professional recruitment tests, validated online resources, and additional items generated with GPT-4o based on authoritative Persian psychology textbooks. It also describes item review for factual accuracy and relevance, removal of incomplete, ambiguous, or low-quality questions, validation of four plausible answers per question, and elimination of copyright-sensitive content [2510.03913].

The common design logic across these descriptions is clear. PsychoLexEval is not presented as a web-scraped benchmark with minimal filtering; it is presented as a curated examination-style resource with explicit expert review and format normalization. This distinguishes it from broader, less specialized evaluation sets and helps explain why it is used as a screening benchmark rather than only as a post hoc leaderboard instrument.

| Aspect | Reported description |
|---|---|
| Format | MCQA with four answer options |
| Size | 3,430 items |
| Sources | Exams, online resources, GPT-generated questions |
| Quality control | Expert review, filtering, copyright screening |

## 4. Evaluation protocol and scoring

The original evaluation protocol tests LLMs in three prompting contexts: zero-shot, one-shot, and five-shot. The primary metric is accuracy, defined as the proportion of correct answers out of total questions. The paper also reports that identical generation configurations are used across models for consistency, and that only open-source LLMs are compared for accessibility and reproducibility [2408.08848].

The later PsychoLexTherapy usage narrows the protocol to model screening in a zero-shot setting. There, models answer MCQs without in-context examples or fine-tuning, the output is constrained to a single choice per question, temperature is set to 0.01, top-p to 0.9, and the output is the most probable single choice. Accuracy is again the core metric, expressed as the percentage of correctly answered questions [2510.03913].

In practice, this means PsychoLexEval supports at least two evaluative regimes. The first is comparative benchmarking across shot settings and languages. The second is threshold-oriented screening for downstream system design. This suggests that the benchmark’s multiple-choice structure is central to its reuse: MCQ formatting enables standardized prompting, deterministic answer extraction, and direct accuracy-based comparison without requiring a secondary judge model.

## 5. Reported empirical findings

The original PsychoLex study reports that model size and domain specialization both matter on PsychoLexEval. Larger models such as Llama-3.1 Instruct 70B generally outperform smaller ones, but domain-specific adaptation also yields strong gains. PsychoLexLLaMA, specialized via pre-training and fine-tuning on psychological content, matches or exceeds general-purpose models in several settings, especially among 8B-scale models. The paper further reports that Persian models show larger gains from additional context than English models when moving from zero-shot to five-shot prompting [2408.08848].

Sample results reported in the original work include, for Persian, average accuracy of 69.52% for Llama-3.1 Instruct 70B, 62.85% for PsychoLexLLaMA-average 70B, and 45.85% for PsychoLexLLaMA-average 8B. For English, the reported averages are 92.58% for Llama-3.1 Instruct 70B, 91.95% for PsychoLexLLaMA-average 70B, and 89.72% for PsychoLexLLaMA-average 8B [2408.08848].

In PsychoLexTherapy, PsychoLexEval is used to compare sub-10B Persian-capable SLMs for on-device feasibility and privacy. Reported accuracies are 55.2 for Gemma-3 7.8B, 53.0 for Qwen-3 8.2B, 50.4 for Gemma-3 4.3B, 48.3 for Qwen-3 4.0B, 33.1 for Gemma-3 1.0B, 31.2 for Mistral 7.2B, 28.7 for LLaMA-3.2 3.2B, and 21.3 for LLaMA-3.2 1.2B. The accompanying interpretation is that higher-parameter SLMs perform best, while smaller models perform considerably worse [2510.03913].

These results support two recurrent conclusions in the reported literature. First, psychological competence as measured by PsychoLexEval is sensitive to both scale and specialization. Second, the benchmark is sufficiently discriminative to shape downstream architectural choices, including privacy-preserving, on-device systems.

## 6. Position within the broader evaluation ecosystem

PsychoLexEval occupies a narrower but more focused position than several adjacent psychological or mental-health benchmarks. Unlike PsyEval, which is organized as a suite spanning knowledge, diagnosis, therapeutic conversation, empathy understanding, and safety understanding, PsychoLexEval is MCQA-centric and primarily oriented toward probing psychological knowledge and reasoning in a controlled format [2311.09189]. Unlike CPsyExam, which separates psychological knowledge from case analysis and includes multiple-choice, multiple-response, and open-ended questions drawn from Chinese examinations, PsychoLexEval is described in the provided sources as a four-option MCQ benchmark [2405.10212].

This makes a common misconception worth addressing directly: PsychoLexEval is not a counseling dialogue benchmark, a psychotherapy simulation benchmark, or a safety benchmark. It is instead a curated examination-style benchmark for knowledge-oriented evaluation. Downstream systems may use it as a prerequisite for therapeutic applications, but the benchmark itself does not evaluate multi-turn counseling quality, empathy, or clinical safety in the same way as dialogue-centered resources do [2510.03913].

At the same time, its narrowness is part of its utility. PsychoLexTherapy explicitly uses PsychoLexEval as a gateway benchmark so that only models above a minimum competence threshold proceed to more complex therapeutic reasoning tasks. This suggests a layered evaluation philosophy: first establish psychological literacy with MCQs, then evaluate dialogue, memory, empathy, coherence, cultural fit, and personalization in subsequent stages [2510.03913].

Within that layered framework, PsychoLexEval functions as an infrastructural benchmark. It provides a standardized, reusable, culturally grounded basis for comparing models in psychology, especially in Persian and English settings, and for distinguishing raw domain knowledge from later-stage therapeutic behavior [2408.08848].

Source: https://www.emergentmind.com/topics/psycholexeval