Establish feature correspondences and nonlinear relations across SAE configurations

Establish whether validated feature spaces learned by different sparse autoencoder configurations admit one-to-one feature correspondences and characterize nonlinear relationships between those feature spaces.

Background

The paper probes whether larger or more active sparse autoencoders produce finer-grained semantic representations by fitting affine linear maps between validated feature-activation spaces from different SAE configurations. The observed directional asymmetry is consistent with larger SAEs containing enough information to reconstruct features from smaller SAEs.

The authors explicitly qualify this result: the probe is only a diagnostic of representational resolution, not evidence that individual features split cleanly across configurations. One-to-one correspondences between features and possible nonlinear relations remain unresolved.

References

This probe does not establish one-to-one feature correspondences, and it may underestimate nonlinear relationships between feature spaces.

Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features  (2609.09575 - Joh et al., 9 Sep 2026) in Appendix, Section “Feature Granularity Across SAE Configurations,” subsection “Interpretation and limitations”