Generalization Beyond Gemma Models and Gemma Scope 2 SAEs

Establish whether SAEVerbalizer generalizes to LLM families, sparse autoencoder architectures, and representation spaces beyond the Gemma LLMs and Gemma Scope 2 SAEs evaluated in the study.

Background

The experiments primarily evaluate SAEVerbalizer using Gemma LLMs and Gemma Scope 2 sparse autoencoders. Although the framework is designed to verbalize decoder directions and includes cross-LLM adapter experiments, the reported evidence does not cover other LLM families, SAE designs, or representation spaces. Consequently, whether the observed verbalization and transfer capabilities persist under broader architectural and representational conditions remains unresolved.

References

We primarily evaluate Gemma LLMs and Gemma Scope 2 SAEs, so generalization to other LLM families, SAE architectures, and representation spaces remains to be established.

— SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization  (2608.13538 - Meng et al., 13 Aug 2026) in Limitations, paragraph “LLM and SAE Coverage”

It remains unclear whether the results generalize to other model families, instruction-tuned models, mixture-of-experts architectures, or SAEs trained on MLP and attention-head outputs rather than the residual stream.

— When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs  (2608.25941 - Gupte et al., 26 Aug 2026) in Section 5, “Limitations and Future Work” (Section \ref{sec:limitations})

Generalization to larger or differently architected models, to other reasoning domains (e.g., commonsense, code), and to truly extremely low-resource languages remains to be verified.

— Enhancing Low-Resource Language Reasoning via High-Resource Language Feature Transfer  (2608.30462 - Song et al., 31 Aug 2026) in Limitations, paragraph “Models and benchmarks”