Identify the properties that predict interpretability

Identify the properties of machine-learning models that predict their interpretability, including whether alignment with human perceptual or semantic organization, feature-activation locality, or robustness explains differences in human understandability.

Background

The paper reports that interpretability is dissociated from both model scale and task performance, so commonly used indicators of capability do not explain why some models are more interpretable than others. It proposes alignment with human organization, activation locality, and robustness as candidate predictors, but states that principled explanations for interpretability remain unresolved.

References

What predicts interpretability is still unresolved.

From Interpretability Methods to Interpretable Models  (2609.05399 - Colin et al., 4 Sep 2026) in Section 4, paragraph “What predicts interpretability is still unresolved”