Define and evaluate translations for many-to-one terminology mappings

Determine how translations involving multiple source synonyms mapped to a single prescribed target should ideally be formulated and evaluated in terminology-aware translation, particularly when faithful adherence requires repetition that conventional quality metrics penalize.

Background

The paper identifies a conflict between translation quality and terminology adherence when a glossary maps several distinct source expressions to one target expression. For example, translating both “boot lid” and “trunk lid” with the same prescribed term may be required for glossary adherence, even though a quality estimator favors replacing one occurrence with a semantically related alternative to avoid repetition.

This unresolved issue affects both the design of terminology-aware translation systems and the evaluation of their outputs. The authors observe that quality estimation may reward edits that damage prescribed-term adherence, indicating that current quality metrics and terminology metrics can have opposing preferences in such cases.

References

What the translations of these sentences should ideally look like, and how they should be scored in the terminology task, remain open questions.

SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better Teachers  (2609.09999 - Liao et al., 9 Sep 2026) in Section Discussion, paragraph “Many-to-one glossaries make quality and adherence disagree.”