Cross-Lingual Consistency of Human DDP-Stratified Preferences

Ascertain whether human preferences for professional human post-edits over automatic post-edits exhibit the same discourse-dependency-stratified pattern across language pairs.

Background

The English–Chinese appendix finds that automatic metrics remain largely insensitive to DDP, but it does not provide the human evaluation needed to determine whether the English–Korean pattern also holds for human judgments. The smaller gap between automatic post-editing and human post-editing in English–Chinese may reflect stronger Chinese coverage by the evaluated models or differences in discourse phenomena such as zero anaphora and topic drop.

Consequently, the paper leaves unresolved whether human evaluators would show the same increasing preference for human revisions as discourse dependency grows when the target language and language-pair characteristics change.

References

Distinguishing the two requires human evaluation, and whether human preferences follow the same DDP-stratified pattern across language pairs remains open.

Discourse Dependency: A Continuous Criterion for Translation Difficulty  (2609.04959 - Kim et al., 4 Sep 2026) in Appendix A.4, “Cross-lingual Consistency: En–Zh”