Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
5 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Global Voices, Local Biases: Socio-Cultural Prejudices across Languages (2310.17586v1)

Published 26 Oct 2023 in cs.CL

Abstract: Human biases are ubiquitous but not uniform: disparities exist across linguistic, cultural, and societal borders. As large amounts of recent literature suggest, LLMs (LMs) trained on human data can reflect and often amplify the effects of these social biases. However, the vast majority of existing studies on bias are heavily skewed towards Western and European languages. In this work, we scale the Word Embedding Association Test (WEAT) to 24 languages, enabling broader studies and yielding interesting findings about LM bias. We additionally enhance this data with culturally relevant information for each language, capturing local contexts on a global scale. Further, to encompass more widely prevalent societal biases, we examine new bias dimensions across toxicity, ableism, and more. Moreover, we delve deeper into the Indian linguistic landscape, conducting a comprehensive regional bias analysis across six prevalent Indian languages. Finally, we highlight the significance of these social biases and the new dimensions through an extensive comparison of embedding methods, reinforcing the need to address them in pursuit of more equitable LLMs. All code, data and results are available here: https://github.com/iamshnoo/weathub.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Anjishnu Mukherjee (6 papers)
  2. Chahat Raj (11 papers)
  3. Ziwei Zhu (59 papers)
  4. Antonios Anastasopoulos (111 papers)
Citations (14)

Summary

We haven't generated a summary for this paper yet.