Generalization of metric reliability to multilingual encoders

Establish whether the reliability ranking of cross-lingual sharing metrics reported for decoder-only language models transfers to multilingual encoder models.

Background

The empirical study evaluates 21 decoder-only LLMs, whereas a substantial portion of prior cross-lingual representation research concerns multilingual encoders. Because encoder and decoder architectures differ in training objectives and attention patterns, the observed robustness of ILO and ANC and the failures associated with GMM dominance and CKA may not generalize across model types.

The paper explicitly leaves unresolved whether its metric-reliability conclusions apply to multilingual encoders.

References

Whether the reliability ranking we report transfers to encoders is untested, and we make no claim about them.

A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures  (2609.04819 - Holmström et al., 4 Sep 2026) in Section ‘Decoder-only models,’ Limitations

Second, we cannot say from these data whether ILO remains the most reliable metric, or whether \GMMLAMBDA\ fails in the same way, for languages thinly represented in pretraining, the regime where a practitioner would most want a trustworthy sharing metric; testing this requires benchmarks with wider coverage and is left to future work.

A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures  (2609.04819 - Holmström et al., 4 Sep 2026) in Section ‘Language selection and coverage,’ Limitations