Generalization Beyond Gemma Models and Gemma Scope 2 SAEs
Establish whether SAEVerbalizer generalizes to LLM families, sparse autoencoder architectures, and representation spaces beyond the Gemma LLMs and Gemma Scope 2 SAEs evaluated in the study.
References
We primarily evaluate Gemma LLMs and Gemma Scope 2 SAEs, so generalization to other LLM families, SAE architectures, and representation spaces remains to be established.
It remains unclear whether the results generalize to other model families, instruction-tuned models, mixture-of-experts architectures, or SAEs trained on MLP and attention-head outputs rather than the residual stream.
Generalization to larger or differently architected models, to other reasoning domains (e.g., commonsense, code), and to truly extremely low-resource languages remains to be verified.