Scaling translation and monolingual fine-tuning analyses across language pairs

Determine whether the observed translation trends and the ability of monolingual fine-tuning to recover the performance of parallel fine-tuning persist across a large number of language pairs, and identify the other factors that explain when monolingual fine-tuning works well and when it does not.

Background

The translation experiments examine a relatively small number of language pairs, with the principal fine-tuning comparison centered on Xhosa and selected additional languages. It therefore remains unresolved how broadly the reported transfer from monolingual fine-tuning to translation applies, and which properties of languages, models, data, or translation directions determine its success.

References

Future work could scale up this experiment to investigate whether these trends hold across a large number of language pairs, and to what degree monolingual fine-tuning can recover the performance of parallel fine-tuning across many language pairs (and what other factors explain when this works well and when it does not).

The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space  (2609.00515 - Brinton et al., 1 Sep 2026) in Section Limitations