Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
41 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
41 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

CultureBank: An Online Community-Driven Knowledge Base Towards Culturally Aware Language Technologies (2404.15238v1)

Published 23 Apr 2024 in cs.CL and cs.AI

Abstract: To enhance LLMs' cultural awareness, we design a generalizable pipeline to construct cultural knowledge bases from different online communities on a massive scale. With the pipeline, we construct CultureBank, a knowledge base built upon users' self-narratives with 12K cultural descriptors sourced from TikTok and 11K from Reddit. Unlike previous cultural knowledge resources, CultureBank contains diverse views on cultural descriptors to allow flexible interpretation of cultural knowledge, and contextualized cultural scenarios to help grounded evaluation. With CultureBank, we evaluate different LLMs' cultural awareness, and identify areas for improvement. We also fine-tune a LLM on CultureBank: experiments show that it achieves better performances on two downstream cultural tasks in a zero-shot setting. Finally, we offer recommendations based on our findings for future culturally aware language technologies. The project page is https://culturebank.github.io . The code and model is at https://github.com/SALT-NLP/CultureBank . The released CultureBank dataset is at https://huggingface.co/datasets/SALT-NLP/CultureBank .

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (8)
  1. Weiyan Shi (41 papers)
  2. Ryan Li (13 papers)
  3. Yutong Zhang (34 papers)
  4. Caleb Ziems (22 papers)
  5. Chunhua yu (2 papers)
  6. Raya Horesh (10 papers)
  7. Rogério Abreu de Paula (2 papers)
  8. Diyi Yang (151 papers)
Citations (17)
Github Logo Streamline Icon: https://streamlinehq.com

GitHub

X Twitter Logo Streamline Icon: https://streamlinehq.com

Tweets