Does Incomplete Syntax Influence Korean Language Model? Focusing on Word Order and Case Markers
Abstract: Syntactic elements, such as word order and case markers, are fundamental in natural language processing. Recent studies show that syntactic information boosts LLM performance and offers clues for people to understand their learning mechanisms. Unlike languages with a fixed word order such as English, Korean allows for varied word sequences, despite its canonical structure, due to case markers that indicate the functions of sentence components. This study explores whether Korean LLMs can accurately capture this flexibility. We note that incomplete word orders and omitted case markers frequently appear in ordinary Korean communication. To investigate this further, we introduce the Syntactically Incomplete Korean (SIKO) dataset. Through SIKO, we assessed Korean LLMs' flexibility with incomplete syntax and confirmed the dataset's training value. Results indicate these models reflect Korean's inherent flexibility, accurately handling incomplete inputs. Moreover, fine-tuning with SIKO enhances the ability to handle common incomplete Korean syntactic forms. The dataset's simple construction process, coupled with significant performance enhancements, solidifies its standing as an effective data augmentation technique.
- Word order does matter and shuffled language models know it. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 6907–6919, 2022.
- Language models are few-shot learners. arXiv preprint, 2020.
- Noam Chomsky. Syntactic structures. Mouton de Gruyter, 2002.
- The pascal recognising textual entailment challenge. In Proceedings of the First International Conference on Machine Learning Challenges: Evaluating Predictive Uncertainty Visual Object Classification, and Recognizing Textual Entailment, MLCW’05, pp. 177–190, Berlin, Heidelberg, 2005. Springer-Verlag. ISBN 3540334270. doi: 10.1007/11736790_9. URL https://doi.org/10.1007/11736790_9.
- Simcse: Simple contrastive learning of sentence embeddings, 2022.
- Aeda: an easier data augmentation technique for text classification. arXiv preprint arXiv:2108.13230, 2021.
- Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
- Korean Grammar Theory, volume 1. hakyounsa, Seoul, 1997.
- Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
- OpenAI. Chatgpt, 2022. URL https://beta.openai.com/signup/. Available at: https://beta.openai.com/signup/.
- Dennis Park. pko-t5: Paust korean t5 for text-to-text unified framework, 5 2022. URL https://github.com/paust-team/pko-t5.
- Klue: Korean language understanding evaluation. arXiv preprint arXiv:2105.09680, 2021.
- Out of order: How important is the sequential order of words in a sentence in natural language understanding tasks? arXiv preprint arXiv:2012.15180, 2020.
- Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http://jmlr.org/papers/v21/20-074.html.
- Unnatural language inference. arXiv preprint arXiv:2101.00010, 2020.
- Masked language modeling and the distributional hypothesis: Order word matters pre-training for little. arXiv preprint arXiv:2104.06644, 2021.
- Ho-min Sohn. Korean language in culture and society. University of Hawaii press, 2005.
- What syntax can contribute in the entailment task. In Machine Learning Challenges Workshop, pp. 205–216. Springer, 2005.
- Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461, 2018.
- Eda: Easy data augmentation techniques for boosting performance on text classification tasks. arXiv preprint arXiv:1901.11196, 2019.
- ESimCSE: Enhanced sample building method for contrastive learning of unsupervised sentence embedding. In Proceedings of the 29th International Conference on Computational Linguistics, pp. 3898–3907, Gyeongju, Republic of Korea, October 2022. International Committee on Computational Linguistics. URL https://aclanthology.org/2022.coling-1.342.
- Unsupervised data augmentation for consistency training. Advances in neural information processing systems, 33:6256–6268, 2020.
- Korean: A comprehensive grammar. Routledge, 2013.
- Gpt3mix: Leveraging large-scale language models for text augmentation. arXiv preprint arXiv:2104.08826, 2021.
- Learning from perturbations: Diverse and informative dialogue generation with inverse adversarial training. arXiv preprint arXiv:2105.15171, 2021.
Paper Prompts
Sign up for free to create and run prompts on this paper.