Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
80 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

NusaCrowd: A Call for Open and Reproducible NLP Research in Indonesian Languages (2207.10524v2)

Published 21 Jul 2022 in cs.CL and cs.AI

Abstract: At the center of the underlying issues that halt Indonesian NLP research advancement, we find data scarcity. Resources in Indonesian languages, especially the local ones, are extremely scarce and underrepresented. Many Indonesian researchers do not publish their dataset. Furthermore, the few public datasets that we have are scattered across different platforms, thus makes performing reproducible and data-centric research in Indonesian NLP even more arduous. Rising to this challenge, we initiate the first Indonesian NLP crowdsourcing effort, NusaCrowd. NusaCrowd strives to provide the largest datasheets aggregation with standardized data loading for NLP tasks in all Indonesian languages. By enabling open and centralized access to Indonesian NLP resources, we hope NusaCrowd can tackle the data scarcity problem hindering NLP progress in Indonesia and bring NLP practitioners to move towards collaboration.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (11)
  1. Samuel Cahyawijaya (75 papers)
  2. Alham Fikri Aji (94 papers)
  3. Holy Lovenia (30 papers)
  4. Genta Indra Winata (94 papers)
  5. Bryan Wilie (24 papers)
  6. Rahmad Mahendra (14 papers)
  7. Fajri Koto (47 papers)
  8. David Moeljadi (5 papers)
  9. Karissa Vincentio (5 papers)
  10. Ade Romadhony (4 papers)
  11. Ayu Purwarianti (39 papers)
Citations (4)