Determine generalizability of mechanistic interpretability findings across model families
Determine the extent to which mechanistic interpretability findings derived from CNN-based image models, BERT-based text models, and GPT-based language models generalize to other architectures and application contexts.
References
The degree of generalizability of these findings to other models and contexts is currently a somewhat open question.
Whether it generalizes to other frontier mixture-of-experts models, and what carries the refusal that resists, remain unmeasured.
The generalizability of the UED to non-mechanical domains, which typically require different model types for investigation, remains to be validated.
Several questions remain open. The middle layers may be linear, as the last layer of Pythia-410m is, but an instrument with finer resolution is needed to tell. The cost gap at the last layer of Pythia-160m, absent in Pythia-410m, may close with scale or may just be an artifact of Pythia's models. Finally, nothing in our results says whether the agreement with optimal transport is a property of LLMs or of trained transformers in general.
Consequently, it remains unexplored whether this distribution-aware neuron identification generalizes effectively to models employing standard MLP architectures that lack this bifurcated signal structure.