Since the Scientific Literature Is Multilingual, Our Models Should Be Too

Published 27 Mar 2024 in cs.CL | (2403.18251v1)

Abstract: English has long been assumed the $\textit{lingua franca}$ of scientific research, and this notion is reflected in the NLP research involving scientific document representation. In this position piece, we quantitatively show that the literature is largely multilingual and argue that current models and benchmarks should reflect this linguistic diversity. We provide evidence that text-based models fail to create meaningful representations for non-English papers and highlight the negative user-facing impacts of using English-only models non-discriminately across a multilingual domain. We end with suggestions for the NLP community on how to improve performance on non-English documents.