Generalization of metric reliability to multilingual encoders
Establish whether the reliability ranking of cross-lingual sharing metrics reported for decoder-only language models transfers to multilingual encoder models.
References
Whether the reliability ranking we report transfers to encoders is untested, and we make no claim about them.
— A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures
(2609.04819 - Holmström et al., 4 Sep 2026) in Section ‘Decoder-only models,’ Limitations
Second, we cannot say from these data whether ILO remains the most reliable metric, or whether \GMMLAMBDA\ fails in the same way, for languages thinly represented in pretraining, the regime where a practitioner would most want a trustworthy sharing metric; testing this requires benchmarks with wider coverage and is left to future work.
— A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures
(2609.04819 - Holmström et al., 4 Sep 2026) in Section ‘Language selection and coverage,’ Limitations