Characterize refusal mechanisms at the representation level

Determine how refusal mechanisms in safety-aligned transformer-based large language models operate in terms of their underlying internal token representations across layers, clarifying the representational implementation of refusal beyond activation-space directions.

Background

Prior work has identified that refusal behavior in aligned LLMs can be mediated by specific directions in activation space. However, the representational basis of refusal—how internal token representations encode and trigger refusal across layers—has not been clearly described. The paper motivates this gap and studies representation hijacking during inference, highlighting the need to understand refusal at the representation level.

References

While recent works identified refusal directions that emerge in the activation space \citep{arditi2024refusal}, it remains unclear how these refusal mechanisms operate in terms of the underlying representations. This is especially crucial since those representations can be changed in-context as the result of user prompts \citep{park2025iclr}.

In-Context Representation Hijacking  (2512.03771 - Yona et al., 3 Dec 2025) in Section 1, Introduction

Mediation: our design establishes causal sufficiency (inducing and suppressing abstention); a full mediation analysis of natural abstention (e.g., activation patching between matched control/void pairs, for which TRAPSBench's minimal-pair structure is well suited) is left to future work.

TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint  (2608.13167 - Pramono et al., 13 Aug 2026) in Appendix, Section “Activation Steering Details,” subsection “Interpretation,” paragraph “Scope of the causal claim”

They do not establish that all visual safety mechanisms are diffuse.

Do VLMs Share Safety Neurons Across Modalities?  (2608.30750 - Li et al., 31 Aug 2026) in Section 5, subsection “Text vs. Visual Safety Neurons,” paragraph “Representation-level evidence indicates a higher-dimensional visual safety subspace”

We leave tracing full reasons behind this behavior of Qwen3-8B for future work.

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness  (2609.03887 - Nguyen et al., 3 Sep 2026) in Appendix, Section Refusal Direction Cosine Similarity