Generalization Beyond Gemma Models and Gemma Scope 2 SAEs

Establish whether SAEVerbalizer generalizes to LLM families, sparse autoencoder architectures, and representation spaces beyond the Gemma LLMs and Gemma Scope 2 SAEs evaluated in the study.

Background

The experiments primarily evaluate SAEVerbalizer using Gemma LLMs and Gemma Scope 2 sparse autoencoders. Although the framework is designed to verbalize decoder directions and includes cross-LLM adapter experiments, the reported evidence does not cover other LLM families, SAE designs, or representation spaces. Consequently, whether the observed verbalization and transfer capabilities persist under broader architectural and representational conditions remains unresolved.

References

We primarily evaluate Gemma LLMs and Gemma Scope 2 SAEs, so generalization to other LLM families, SAE architectures, and representation spaces remains to be established.

SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization  (2608.13538 - Meng et al., 13 Aug 2026) in Limitations, paragraph “LLM and SAE Coverage”