Explain the mechanism by which hard examples improve glossary adherence

Investigate whether the improvement produced by training on hard terminology examples is mechanistically attributable to teaching models to consult the glossary when their default translation differs from the prescribed term, and identify what model mechanisms implement this behavior.

Background

The paper’s controlled study shows that increasing the proportion of hard examples—examples for which SalamandraTA-7b-instruct v2.0 misses at least one prescribed term—substantially improves term accuracy at fixed data volume. The authors interpret this result as evidence that hard examples teach a behavioral pattern: consulting and following the glossary when the model’s unconstrained translation would otherwise choose a different term.

The paper does not establish whether this behavioral interpretation is mechanistically correct or explain which internal model processes implement it. Clarifying these points could guide more principled data-selection methods and determine whether the observed gains arise from the proposed glossary-reading behavior or another mechanism.

References

Whether this behavioural account holds up mechanistically, and what in the model implements it, we leave to future work.

SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better Teachers  (2609.09999 - Liao et al., 9 Sep 2026) in Section Discussion, paragraph “What the filter buys.”