Characterize refusal mechanisms at the representation level
Determine how refusal mechanisms in safety-aligned transformer-based large language models operate in terms of their underlying internal token representations across layers, clarifying the representational implementation of refusal beyond activation-space directions.
References
While recent works identified refusal directions that emerge in the activation space \citep{arditi2024refusal}, it remains unclear how these refusal mechanisms operate in terms of the underlying representations. This is especially crucial since those representations can be changed in-context as the result of user prompts \citep{park2025iclr}.
Mediation: our design establishes causal sufficiency (inducing and suppressing abstention); a full mediation analysis of natural abstention (e.g., activation patching between matched control/void pairs, for which TRAPSBench's minimal-pair structure is well suited) is left to future work.
They do not establish that all visual safety mechanisms are diffuse.
We leave tracing full reasons behind this behavior of Qwen3-8B for future work.