Resolve the context-access confound in interpreter-system comparisons
Determine whether the performance gap between the MT-specific systems and LLM interpreters, and the observed layer-wise degradation pattern in semantic, pragmatic, and cultural-social communicative success, arise from architectural capability differences or from the unequal context provided to the systems.
References
We do not run a matched condition that gives the MT systems comparable context or restricts the LLMs to raw source text, so this confound remains open.
— Evaluating Communicative Success in Machine-Translated Conversation
(2609.19885 - Haznitrama et al., 17 Sep 2026) in Appendix, Section Limitations, subsection “Context asymmetry”