Program Translation via Code Distillation (2310.11476v1)

Published 17 Oct 2023 in cs.SE and cs.LG

Abstract: Software version migration and program translation are an important and costly part of the lifecycle of large codebases. Traditional machine translation relies on parallel corpora for supervised translation, which is not feasible for program translation due to a dearth of aligned data. Recent unsupervised neural machine translation techniques have overcome data limitations by included techniques such as back translation and low level compiler intermediate representations (IR). These methods face significant challenges due to the noise in code snippet alignment and the diversity of IRs respectively. In this paper we propose a novel model called Code Distillation (CoDist) whereby we capture the semantic and structural equivalence of code in a language agnostic intermediate representation. Distilled code serves as a translation pivot for any programming language, leading by construction to parallel corpora which scale to all available source code by simply applying the distillation compiler. We demonstrate that our approach achieves state-of-the-art performance on CodeXGLUE and TransCoder GeeksForGeeks translation benchmarks, with an average absolute increase of 12.7% on the TransCoder GeeksforGeeks translation benchmark compare to TransCoder-ST.

PDF HTML Abstract

Summarize Bookmark Chat (Pro)

Authors (7)

Yufan Huang (20 papers)
Mengnan Qi (5 papers)
Yongqiang Yao (21 papers)
Maoquan Wang (7 papers)
Bin Gu (86 papers)
Colin Clement (10 papers)
Neel Sundaresan (38 papers)

Citations (3)

View on Semantic Scholar

Program Translation via Code Distillation (2310.11476v1)

Related Papers