Establish feature correspondences and nonlinear relations across SAE configurations
Establish whether validated feature spaces learned by different sparse autoencoder configurations admit one-to-one feature correspondences and characterize nonlinear relationships between those feature spaces.
References
This probe does not establish one-to-one feature correspondences, and it may underestimate nonlinear relationships between feature spaces.
— Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features
(2609.09575 - Joh et al., 9 Sep 2026) in Appendix, Section “Feature Granularity Across SAE Configurations,” subsection “Interpretation and limitations”