Evaluating and Mitigating Anti-LGBTQ Biases in German and Multilingual Language Models
Abstract: While gender and racial biases in LLMs have been widely studied, anti-LGBTQ biases remain underexplored, particularly beyond English. Existing benchmarks often do not capture cultural and linguistic variation and rely on gender representations. This paper introduces a multilingual German-English benchmark dataset for the evaluation of anti-LGBTQ biases in LLMs. It combines community-sourced stereotypes from German-speaking queer individuals with a German translation of WinoQueer. The data is used to evaluate eight LLMs across sizes and architectures and explore mitigation through fine-tuning on community and progressive media content. Results show that LLMs reproduce anti-queer stereotypes, with variation across identities and models. Differences between the translated and community-based data highlight the importance of cultural adaptation for multilingual bias evaluation. Fine-tuning reduces bias on average, but not consistently across models and identities. Warning: This text contains examples of anti-queer hateful language and stereotypes.
Paper Prompts
Sign up for free to create and run prompts on this paper.