Annotating Norwegian Language Varieties on Twitter for Part-of-Speech (2210.06150v1)

Published 12 Oct 2022 in cs.CL

Abstract: Norwegian Twitter data poses an interesting challenge for NLP tasks. These texts are difficult for models trained on standardized text in one of the two Norwegian written forms (Bokm{\aa}l and Nynorsk), as they contain both the typical variation of social media text, as well as a large amount of dialectal variety. In this paper we present a novel Norwegian Twitter dataset annotated with POS-tags. We show that models trained on Universal Dependency (UD) data perform worse when evaluated against this dataset, and that models trained on Bokm{\aa}l generally perform better than those trained on Nynorsk. We also see that performance on dialectal tweets is comparable to the written standards for some models. Finally we perform a detailed analysis of the errors that models commonly make on this data.

Citations (5)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Annotating Norwegian Language Varieties on Twitter for Part-of-Speech (2210.06150v1)

Summary

Related Papers