Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
51 tokens/sec
GPT-4o
60 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
8 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

PsyEval: A Suite of Mental Health Related Tasks for Evaluating Large Language Models (2311.09189v2)

Published 15 Nov 2023 in cs.CL

Abstract: Evaluating LLMs in the mental health domain poses distinct challenged from other domains, given the subtle and highly subjective nature of symptoms that exhibit significant variability among individuals. This paper presents PsyEval, the first comprehensive suite of mental health-related tasks for evaluating LLMs. PsyEval encompasses five sub-tasks that evaluate three critical dimensions of mental health. This comprehensive framework is designed to thoroughly assess the unique challenges and intricacies of mental health-related tasks, making PsyEval a highly specialized and valuable tool for evaluating LLM performance in this domain. We evaluate twelve advanced LLMs using PsyEval. Experiment results not only demonstrate significant room for improvement in current LLMs concerning mental health but also unveil potential directions for future model optimization.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (6)
  1. Haoan Jin (2 papers)
  2. Siyuan Chen (92 papers)
  3. Mengyue Wu (57 papers)
  4. Kenny Q. Zhu (50 papers)
  5. Dilawaier Dilixiati (1 paper)
  6. Yewei Jiang (1 paper)
Citations (3)