Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
102 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Learn to Code-Switch: Data Augmentation using Copy Mechanism on Language Modeling (1810.10254v2)

Published 24 Oct 2018 in cs.CL

Abstract: Building large-scale datasets for training code-switching LLMs is challenging and very expensive. To alleviate this problem using parallel corpus has been a major workaround. However, existing solutions use linguistic constraints which may not capture the real data distribution. In this work, we propose a novel method for learning how to generate code-switching sentences from parallel corpora. Our model uses a Seq2Seq model in combination with pointer networks to align and choose words from the monolingual sentences and form a grammatical code-switching sentence. In our experiment, we show that by training a LLM using the augmented sentences we improve the perplexity score by 10% compared to the LSTM baseline.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Genta Indra Winata (94 papers)
  2. Andrea Madotto (65 papers)
  3. Chien-Sheng Wu (77 papers)
  4. Pascale Fung (151 papers)
Citations (19)