Universality of Features and Circuits

Determine the degrees of universality of features and circuits across transformer-based language models and tasks, and ascertain how this universality depends on model training factors such as random initialization, model size, and the loss function used during training.

Background

Mechanistic interpretability investigates whether similar features and circuits recur across different LLMs and tasks, a property referred to as universality. Some studies report recurring components (e.g., induction heads, successor heads), while other work shows qualitatively different circuits emerging under different initializations or low rates of universal neurons across GPT-2 models.

Because many mechanistic analyses have been performed on toy or small models, establishing universality would enable transferring insights to larger models with less bespoke effort. Mixed empirical findings highlight the need to rigorously characterize the extent and conditions under which universality holds.

References

Understanding the degrees of feature and circuit universality and their dependency on various aspects of model training (e.g., initialization, model size, and loss function) remains a crucial open problem.

A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models  (2407.02646 - Rai et al., 2024) in Findings and Applications, Findings on Universality (Section 7, Subsection "Findings on Universality")

However, cross-model representational alignment---specifically whether distinct architectures utilize shared value directions---remains unresolved; the cross-model activation patching and representational comparisons below address this question on the subset of disagreement items.

GYROval: A Robust Benchmark for Cultural Value Orientation in Large Language Models  (2609.18384 - Didenko et al., 16 Sep 2026) in Section 2.10, “Mechanistic sub-study”

Untested is whether a different model refit, input distribution, graph granularity, or pruning procedure returns the same circuit; each is known to change the answer \citep{bali2026headstability,makou2026manycircuits,parekh2026circus}.

Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit  (2608.27254 - Kumar, 27 Aug 2026) in Paragraph 'Three seeds test optimization stability, not uniqueness,' Section 'Limitations in Detail' (Appendix~\ref{app:limitations})

Generalization of the two-stage circuit structure to these settings, and its stability under prompt variation, remains an empirical question.

Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation  (2609.01604 - Vasava et al., 1 Sep 2026) in Section 5, paragraph “Model and task scope”

Therefore generalizability of our findings to models with larger scales remains an open question, given that circuit analysis literature about scale consistency tops out around 2.8B with mixed results \citep{tigges2024llm}.

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness  (2609.03887 - Nguyen et al., 3 Sep 2026) in Limitations section

As such, there is still uncertainty about whether the circuit is fully unique to coreference.

A Circuit for Plural Reference: How LLMs Represent and Retrieve Singular and Plural Entities  (2609.03687 - Danh et al., 3 Sep 2026) in Limitations section

However, most existing work on VLMs still focuses on image-based tasks , leaving open questions about how modality-specific circuits emerge and interact in more complex multimodal scenarios such as videos.

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making  (2609.05149 - Testa et al., 4 Sep 2026) in Section 2, Related Work

This paper presents an empirical finding, but at the moment, we do not have a theory that can explain the success of our approach. We can show that circuit overlap and behavioral generalization go together across formats, models, and items, but we cannot yet say why, or under which conditions the relationship should hold or break.

Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning  (2609.04463 - Varda et al., 3 Sep 2026) in Section “Limitations”