General2Specialized LLMs Translation for E-commerce (2403.03689v2)

Published 6 Mar 2024 in cs.CL and cs.AI

Abstract: Existing Neural Machine Translation (NMT) models mainly handle translation in the general domain, while overlooking domains with special writing formulas, such as e-commerce and legal documents. Taking e-commerce as an example, the texts usually include amounts of domain-related words and have more grammar problems, which leads to inferior performances of current NMT methods. To address these problems, we collect two domain-related resources, including a set of term pairs (aligned Chinese-English bilingual terms) and a parallel corpus annotated for the e-commerce domain. Furthermore, we propose a two-step fine-tuning paradigm (named G2ST) with self-contrastive semantic enhancement to transfer one general NMT model to the specialized NMT model for e-commerce. The paradigm can be used for the NMT models based on LLMs. Extensive evaluations on real e-commerce titles demonstrate the superior translation quality and robustness of our G2ST approach, as compared with state-of-the-art NMT models such as LLaMA, Qwen, GPT-3.5, and even GPT-4.

PDF HTML Abstract

Summarize Bookmark Chat (Pro)

References (19)

Authors (9)

Kaidi Chen (2 papers)
Ben Chen (23 papers)
Dehong Gao (26 papers)
Huangyu Dai (4 papers)
Wen Jiang (52 papers)
Wei Ning (48 papers)
Shanqing Yu (41 papers)
Libin Yang (17 papers)
Xiaoyan Cai (15 papers)

Citations (4)

View on Semantic Scholar

Tweets

https://twitter.com/gm8xx8/status/1765569230219448564

General2Specialized LLMs Translation for E-commerce (2403.03689v2)

Related Papers

Tweets