Generalization of SRT across broader MLLM architectures

Determine whether Safety-awareness Representation Transfer (SRT) generalizes reliably to multimodal large language models beyond the evaluated model families, scales, and training paradigms, given that different models may encode refusal-related computation in different layers or rely on different internal safety mechanisms.

Background

Safety-awareness Representation Transfer (SRT) is evaluated on several multimodal LLM families, including Qwen3-VL, Gemma-3, InternVL2, and LLaVA-OneVision, with additional experiments on larger Qwen3-VL and Gemma-3 models. The authors nevertheless caution that these experiments do not establish generalization across the full range of multimodal architectures, model scales, or training paradigms.

The unresolved issue is whether SRT remains stable when applied to models whose refusal-related computations are represented in different layers or implemented through different internal safety mechanisms. Establishing this broader generalization would clarify the method's applicability beyond the architectures directly tested in the paper.

References

Although we evaluate SRT on multiple representative MLLM families, its generalization to broader architectures remains uncertain. Different models may encode refusal-related computation in different layers or rely on different internal safety mechanisms. Further experiments on more model families, scales, and training paradigms are needed to verify the stability of SRT.

Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models  (2609.02082 - Xiao et al., 2 Sep 2026) in Section Limitations