Syntax-aware Data Augmentation for Neural Machine Translation (2004.14200v1)

Published 29 Apr 2020 in cs.CL

Abstract: Data augmentation is an effective performance enhancement in neural machine translation (NMT) by generating additional bilingual data. In this paper, we propose a novel data augmentation enhancement strategy for neural machine translation. Different from existing data augmentation methods which simply choose words with the same probability across different sentences for modification, we set sentence-specific probability for word selection by considering their roles in sentence. We use dependency parse tree of input sentence as an effective clue to determine selecting probability for every words in each sentence. Our proposed method is evaluated on WMT14 English-to-German dataset and IWSLT14 German-to-English dataset. The result of extensive experiments show our proposed syntax-aware data augmentation method may effectively boost existing sentence-independent methods for significant translation performance improvement.

PDF Abstract

Summarize PDF Markdown Bookmark Chat (Pro)

Authors (4)

Sufeng Duan (13 papers)
Hai Zhao (227 papers)
Dongdong Zhang (79 papers)
Rui Wang (996 papers)

Citations (15)

View on Semantic Scholar

Syntax-aware Data Augmentation for Neural Machine Translation (2004.14200v1)

Related Papers