Resolve the context-access confound in interpreter-system comparisons

Determine whether the performance gap between the MT-specific systems and LLM interpreters, and the observed layer-wise degradation pattern in semantic, pragmatic, and cultural-social communicative success, arise from architectural capability differences or from the unequal context provided to the systems.

Background

The evaluation compares Google Translate, NLLB-200, and SeamlessM4T v2 using only raw source text, whereas the LLM-based interpreter systems receive broader information, including conversation context and cultural-context instructions. This asymmetry is especially consequential because the pragmatic and cultural-social checklist layers evaluate tone, register, honorifics, and related properties that are difficult to infer without contextual information.

The paper therefore leaves unresolved how much of the apparent superiority of LLM interpreters, and how much of the decline from semantic to pragmatic and cultural-social success, should be attributed to model architecture or capability rather than to differences in context access. A matched experiment should either provide comparable context to the MT systems or restrict the LLM interpreters to raw source text.

References

We do not run a matched condition that gives the MT systems comparable context or restricts the LLMs to raw source text, so this confound remains open.

— Evaluating Communicative Success in Machine-Translated Conversation  (2609.19885 - Haznitrama et al., 17 Sep 2026) in Appendix, Section Limitations, subsection “Context asymmetry”