Generalization of SRT across broader MLLM architectures
Determine whether Safety-awareness Representation Transfer (SRT) generalizes reliably to multimodal large language models beyond the evaluated model families, scales, and training paradigms, given that different models may encode refusal-related computation in different layers or rely on different internal safety mechanisms.
References
Although we evaluate SRT on multiple representative MLLM families, its generalization to broader architectures remains uncertain. Different models may encode refusal-related computation in different layers or rely on different internal safety mechanisms. Further experiments on more model families, scales, and training paradigms are needed to verify the stability of SRT.
— Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models
(2609.02082 - Xiao et al., 2 Sep 2026) in Section Limitations