Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
102 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Automatic Extraction of Personality from Text: Challenges and Opportunities (1910.09916v1)

Published 22 Oct 2019 in cs.CL

Abstract: In this study, we examined the possibility to extract personality traits from a text. We created an extensive dataset by having experts annotate personality traits in a large number of texts from multiple online sources. From these annotated texts, we selected a sample and made further annotations ending up in a large low-reliability dataset and a small high-reliability dataset. We then used the two datasets to train and test several machine learning models to extract personality from text, including a LLM. Finally, we evaluated our best models in the wild, on datasets from different domains. Our results show that the models based on the small high-reliability dataset performed better (in terms of $\textrm{R}2$) than models based on large low-reliability dataset. Also, LLM based on small high-reliability dataset performed better than the random baseline. Finally, and more importantly, the results showed our best model did not perform better than the random baseline when tested in the wild. Taken together, our results show that determining personality traits from a text remains a challenge and that no firm conclusions can be made on model performance before testing in the wild.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (5)
  1. Nazar Akrami (2 papers)
  2. Johan Fernquist (1 paper)
  3. Tim Isbister (8 papers)
  4. Lisa Kaati (2 papers)
  5. Björn Pelzer (1 paper)
Citations (10)