Universality of Features and Circuits
Determine the degrees of universality of features and circuits across transformer-based language models and tasks, and ascertain how this universality depends on model training factors such as random initialization, model size, and the loss function used during training.
References
Understanding the degrees of feature and circuit universality and their dependency on various aspects of model training (e.g., initialization, model size, and loss function) remains a crucial open problem.
However, cross-model representational alignment---specifically whether distinct architectures utilize shared value directions---remains unresolved; the cross-model activation patching and representational comparisons below address this question on the subset of disagreement items.
Untested is whether a different model refit, input distribution, graph granularity, or pruning procedure returns the same circuit; each is known to change the answer \citep{bali2026headstability,makou2026manycircuits,parekh2026circus}.
Generalization of the two-stage circuit structure to these settings, and its stability under prompt variation, remains an empirical question.
Therefore generalizability of our findings to models with larger scales remains an open question, given that circuit analysis literature about scale consistency tops out around 2.8B with mixed results \citep{tigges2024llm}.
As such, there is still uncertainty about whether the circuit is fully unique to coreference.
However, most existing work on VLMs still focuses on image-based tasks , leaving open questions about how modality-specific circuits emerge and interact in more complex multimodal scenarios such as videos.
This paper presents an empirical finding, but at the moment, we do not have a theory that can explain the success of our approach. We can show that circuit overlap and behavioral generalization go together across formats, models, and items, but we cannot yet say why, or under which conditions the relationship should hold or break.