Papers
Topics
Authors
Recent
Search
2000 character limit reached

DOSA: A Dataset of Social Artifacts from Different Indian Geographical Subcultures

Published 23 Feb 2024 in cs.CY and cs.CL | (2403.14651v1)

Abstract: Generative models are increasingly being used in various applications, such as text generation, commonsense reasoning, and question-answering. To be effective globally, these models must be aware of and account for local socio-cultural contexts, making it necessary to have benchmarks to evaluate the models for their cultural familiarity. Since the training data for LLMs is web-based and the Web is limited in its representation of information, it does not capture knowledge present within communities that are not on the Web. Thus, these models exacerbate the inequities, semantic misalignment, and stereotypes from the Web. There has been a growing call for community-centered participatory research methods in NLP. In this work, we respond to this call by using participatory research methods to introduce DOSA\textit{DOSA}, the first community-generated D\textbf{D}ataset o\textbf{o}f 615 S\textbf{S}ocial A\textbf{A}rtifacts, by engaging with 260 participants from 19 different Indian geographic subcultures. We use a gamified framework that relies on collective sensemaking to collect the names and descriptions of these artifacts such that the descriptions semantically align with the shared sensibilities of the individuals from those cultures. Next, we benchmark four popular LLMs and find that they show significant variation across regional sub-cultures in their ability to infer the artifacts.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (45)
  1. NITI Aayog. 2023. National multidimensional poverty index.
  2. Large language models associate muslims with violence. Nature Machine Intelligence, 3(6):461–463.
  3. Towards an atlas of cultural commonsense for machine reasoning. arXiv preprint arXiv:2009.05664.
  4. Challenges in designing input method editors for indian lan-guages: The role of word-origin and context. In Proceedings of the Workshop on Advances in Text Input Methods (WTIM 2011), pages 1–9.
  5. Falcon-40B: an open large language model with state-of-the-art performance.
  6. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  7. Probing pre-trained language models for cross-cultural differences in values. arXiv preprint arXiv:2203.13722.
  8. Which humans?
  9. Ready player one! eliciting diverse knowledge using a configurable game. In Proceedings of the ACM Web Conference 2022, pages 1709–1719.
  10. Emily M Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6:587–604.
  11. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  12. Ethical dilemmas, mental health, artificial intelligence, and llm-based chatbots. In International Work-Conference on Bioinformatics and Biomedical Engineering, pages 313–326. Springer.
  13. Assessing cross-cultural alignment between chatgpt and human societies: An empirical study. arXiv preprint arXiv:2303.17466.
  14. Harrison Chase. 2022. LangChain.
  15. The six conundrums of building and deploying language technologies for social good. In ACM SIGCAS/SIGCHI Conference on Computing and Sustainable Societies (COMPASS), pages 12–19.
  16. Bias of ai-generated content: An examination of news produced by large language models. arXiv preprint arXiv:2309.09825.
  17. What does chatgpt return about human values? exploring value bias in chatgpt using a descriptive value theory. arXiv preprint arXiv:2304.03612.
  18. Moral foundations theory: The pragmatic validity of moral pluralism. In Advances in experimental social psychology, volume 47, pages 55–130. Elsevier.
  19. Mining hindi-english transliteration pairs from online hindi lyrics. In LREC, pages 2459–2465.
  20. Speaking multiple languages affects the moral bias of language models. arXiv preprint arXiv:2211.07733.
  21. Challenges and strategies in cross-cultural nlp. arXiv preprint arXiv:2203.10020.
  22. Geert Hofstede. 2011. Dimensionalizing cultures: The hofstede model in context. Online readings in psychology and culture, 2(1):8.
  23. World values survey: Round six - country-pooled datafile. madrid, spain & vienna, austria: Jd systems institute & wvsa secretariat. doi.org/10.14281/18241.8.
  24. Understanding the benefits and challenges of deploying conversational ai leveraging large language models for public health intervention. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–16.
  25. Your spouse needs professional help: Determining the contextual appropriateness of messages through modeling social relationships. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10994–11013.
  26. Gender bias and stereotypes in large language models. arXiv preprint arXiv:2308.14921.
  27. Virginie Mamadouh. 2020. Writing the world in 301 languages: A political geography of the online encyclopedia wikipedia. Handbook of the Changing World Language Map, pages 3801–3824.
  28. More human than human: Measuring chatgpt political bias. Public Choice, pages 1–21.
  29. Extracting cultural commonsense knowledge at scale. In Proceedings of the ACM Web Conference 2023, pages 1907–1917.
  30. OpenAI. 2023. Gpt-4 technical report.
  31. Shramay Palta and Rachel Rudinger. 2023. Fork: A bite-sized test set for probing culinary cultural biases in commonsense reasoning models. In Findings of the Association for Computational Linguistics: ACL 2023, pages 9952–9962.
  32. Cultural incongruencies in artificial intelligence. arXiv preprint arXiv:2211.13069.
  33. Ai’s regimes of representation: A community-centered study of text-to-image models in south asia. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 506–517.
  34. Personality traits in large language models. arXiv preprint arXiv:2307.00184.
  35. Large pre-trained language models contain human-like biases of what is right and wrong to do. Nature Machine Intelligence, 4(3):258–268.
  36. Cultural differences in friendship network behaviors: A snapchat case study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–14.
  37. Devinder Pal Singh and Manoj K Sharma. 2009. Unfolding the indian cultural mosaic: a cross-cultural study of four regional cultures. International Journal of Indian Culture and Business Management, 2(3):247–267.
  38. Janet Stephenson. 2023. Culture and Sustainability: Exploring Stability and Transformation with the Cultures Framework. Springer Nature.
  39. Understanding the capabilities, limitations, and societal impact of large language models. arXiv preprint arXiv:2102.02503.
  40. Exploring large language models’ cognitive moral development through defining issues test. arXiv preprint arXiv:2309.13356.
  41. Vishesh Thakur. 2023. Unveiling gender bias in terms of profession across llms: Analyzing and addressing sociological implications. arXiv preprint arXiv:2307.09162.
  42. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  43. Verbosity: a game for collecting common-sense facts. In Proceedings of the SIGCHI conference on Human Factors in computing systems, pages 75–78.
  44. Evaluating gpt-3 generated explanations for hateful content moderation. arXiv preprint arXiv:2305.17680.
  45. From instructions to intrinsic human values–a survey of alignment goals for big models. arXiv preprint arXiv:2308.12014.
Citations (7)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 5 tweets with 11 likes about this paper.