Mechanistic basis of cross-family perturbation differences

Determine the mechanistic basis for the divergent responses of GPT-2 and Qwen2.5 models to different perturbation types by isolating the effects of grouped-query attention, vocabulary size, and rotary positional embeddings.

Background

The study finds that GPT-2 and Qwen2.5 models respond differently to several perturbation types, with cross-family differences depending on the behavioral or representational metric used. Although the paper identifies grouped-query attention, vocabulary size, and rotary positional embeddings as plausible architectural factors, it does not isolate their individual contributions. Establishing these causal mechanisms would clarify why perturbation signatures vary across model families.

References

Several open questions follow directly from work. First, GPT-2 and Qwen2.5 models exhibit diverging behavior under different perturbation types, but no mechanistic basis for this behavior has been identified in this study. Isolating architectural factors, such as GQA, vocabular size, or RoPE, may provide additional insights into the mechanistic roles played by these model elements.

How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models  (2609.03322 - Chan et al., 3 Sep 2026) in Conclusion