Papers
Topics
Authors
Recent
Search
2000 character limit reached

Refusal Training: Definition, Approaches, and Evaluation

Updated 1 September 2026
  • Refusal training is the set of post-training methods used to make language and multimodal language models decline harmful, unsupported, unknowable, or otherwise inappropriate requests while answering legitimate requests accurately. These methods include supervised fine-tuning, preference optimization, reinforcement fine-tuning, adversarial training, and representation-level interventions.
  • Refusal training addresses two types of refusals: should-not (due to legality, privacy, or safety) and cannot (due to limitations in knowledge, context, or modality). A detailed taxonomy identifies 992 leaf nodes, including categories such as unsupported modalities, missing information, knowledge cutoff, and invalid premises. Effective refusal training ensures that a model should refuse for the right reasons and maintain refusal throughout the entire generation process, avoiding harmful continuations after initial refusals.
  • Effective refusal training must evaluate the full response trajectory to ensure models do not provide harmful information after an initial refusal and must continuously adapt to prevent refusal position bias or compliance with harmful continuations. This involves refining where refusals occur, enforcing persistent refusal during the entire response generation, and balancing false refusals against missed refusals to avoid over-refusal and ensure that harmful compliance responses are minimized, as seen in the DeRTa, NOICE, CRaFT, and refusal token methods.

Refusal training is the set of post-training methods used to make language and multimodal LLMs decline harmful, unsupported, unknowable, or otherwise inappropriate requests while answering legitimate requests accurately. It encompasses supervised fine-tuning, preference optimization, reinforcement finetuning, adversarial training, representation-level interventions, refusal calibration, and evaluation of selective abstention. Contemporary research treats refusal not as a scalar capability to maximize, but as a selective behavioral policy: a reliable model should refuse for the right reasons, preserve helpfulness on safe inputs, avoid hallucinated answers when information is insufficient, and maintain refusal throughout an entire generation rather than only at its beginning.

1. Conceptual foundations and objectives

Refusal training originally focused primarily on harmful prompts: the model receives an unsafe request and a refusal response such as “I cannot assist with that.” The objective is to reduce harmful compliance. Subsequent work has expanded the target behavior to include epistemic and capability limitations, including missing information, false premises, unsupported modalities, knowledge-cutoff questions, and questions whose answers cannot be established from the available evidence.

A central distinction is between should-not and cannot refusals. Should-not refusals arise when the model is capable of carrying out a request but should not do so because of legality, privacy, information hazards, explicit material, or safety policy. Cannot refusals arise from limitations of knowledge, context, modality, skill, identity, or input information. The taxonomy proposed in “Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs” (Recum et al., 2024) identifies 16 top-level labels, including Chain of Command, Legal Compliance / Illegal Activity, Information Hazards, Privacy, NSFW Content, Modalities, Skills / Skill Level, Invalid Premise, Knowledge Cutoff, Unknown Information, Training Data Limits, Missing Context, Missing Identity, Not a Refusal, and Unclear. The detailed taxonomy contains 992 leaf nodes, although the synthetic experiments use 13 categories and do not provide an unambiguous mapping from all 16 top-level labels.

Refusal quality involves at least three distinct dimensions:

  1. Detection: deciding whether to answer or refuse.
  2. Categorization: identifying why refusal is warranted.
  3. Calibration: balancing false refusals against missed refusals.

RefusalBench formalizes this distinction for retrieval-augmented generation as the difference between refusal detection and refusal categorization. Its hierarchical refusal score is defined as:

H=F1×Acat,H = F_1 \times A_{\text{cat}},

where F1F_1 measures binary refusal detection and AcatA_{\text{cat}} measures category accuracy conditional on refusal (Muhamed et al., 12 Oct 2025). This formulation reflects the fact that a model can correctly detect that a context is unreliable while misunderstanding whether the problem is ambiguity, contradiction, missing information, a false premise, granularity mismatch, or epistemic mismatch.

For multimodal systems, “Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal” (Wang et al., 2024) distinguishes extrinsic and intrinsic information boundaries. An extrinsic boundary is exceeded when the image does not contain the information needed to answer. An intrinsic boundary is exceeded when the model lacks the capability, knowledge, perceptual reliability, or confidence required to use available information. The desired behavior is to answer when both boundaries are satisfied and refuse when either is violated.

This view motivates a user-centric value function:

v(i,q,r)={1if r is correct, 0if r is a refusal, −1if r is incorrect.v(i,q,r)= \begin{cases} 1 & \text{if } r \text{ is correct},\ 0 & \text{if } r \text{ is a refusal},\ -1 & \text{if } r \text{ is incorrect}. \end{cases}

The resulting trustworthiness score is:

Strust=2⋅Acc+RefR−1.S_{\text{trust}}=2\cdot\text{Acc}+\text{RefR}-1.

A correct answer is preferred to a refusal, but a refusal is preferred to an incorrect confident answer. This differs from objectives that simply maximize refusal rate.

2. Position, persistence, and composition of refusals

Refusal position bias

“Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training” (Yuan et al., 2024) identifies refusal position bias in conventional safety tuning. Typical refusal examples begin with tokens such as “Sorry,” “I cannot,” or “I apologize,” causing the model to associate refusal with the first few response positions. The paper reports that 99.5% of refusal responses generated by LLaMA3-8B-Instruct and LLaMA3-70B-Instruct place refusal tokens within the first five output tokens.

This creates two weaknesses. First, the model must make an irreversible safety decision before generating additional response content. Second, ordinary safety tuning does not necessarily teach the model how to interrupt a harmful continuation after generation has begun. A model can therefore produce an initial refusal followed by a transition such as “But now that we’ve got that mandatory warning out of the way…” and then provide actionable harmful content.

DeRTa, or Decoupled Refusal Training, addresses this problem with two components:

  • MLE with a Harmful Response Prefix: a randomly selected prefix of a harmful response is appended to a harmful query, and the target is a safe response.
  • Reinforced Transition Optimization (RTO): the model is trained to assign high probability to a refusal token after every harmful prefix.

For a harmful response r^\hat r, RTO optimizes:

LRTO(θ)=−E(q,r^)[∑t=1∣r^∣log⁡Pθ(sorry∣q,r^<t)].\mathcal{L}_{\mathrm{RTO}}(\theta) = -\mathbb{E}_{(q,\hat r)} \left[ \sum_{t=1}^{|\hat r|} \log P_\theta(\textit{sorry}\mid q,\hat r_{<t}) \right].

Despite its name, RTO is an auxiliary likelihood-based term rather than a conventional policy-gradient reinforcement-learning algorithm. MLE with a harmful prefix supplies approximately one harmful-to-safe transition per example, whereas RTO supplies a transition target at every harmful-response position.

On LLaMA3-70B, the combined DeRTa method reduced average ASR from 70.6% to 8.8% across six attacks. On LLaMA3-70B-Instruct, average ASR decreased from 34.9% to approximately 2.2%. The method also shifted the distribution of refusal positions: ordinary safety tuning placed more than 80% of refusals at the beginning, whereas RTO caused 22.3% of responses to contain refusals beyond position 30.

Refusal persistence

“No, of course I can! Refusal Mechanisms Can Be Exploited Using Harmless Fine-Tuning Data” (Kazdan et al., 26 Feb 2025) demonstrates that an initial refusal is not sufficient evidence of safety. It introduces NOICE, a “refuse-then-comply” fine-tuning attack that trains a model to produce a refusal or safety warning, followed by a transition phrase and an answer. The conceptual sequence is:

refusal→safety transition→answer.\text{refusal} \rightarrow \text{safety transition} \rightarrow \text{answer}.

The paper decomposes harmful-response probability as:

P(HR∣HP)=P(HR∣R,HP)P(R∣HP)+P(HR∣¬R,HP)P(¬R∣HP).\mathbb{P}(HR\mid HP) = \mathbb{P}(HR\mid R,HP)\mathbb{P}(R\mid HP) + \mathbb{P}(HR\mid \neg R,HP)\mathbb{P}(\neg R\mid HP).

Conventional shallow attacks primarily increase P(¬R∣HP)\mathbb{P}(\neg R\mid HP) by inducing an agreement prefix. NOICE targets F1F_10 by preserving the refusal prefix while making refusal a preamble to harmful compliance.

With 5,000 training examples, NOICE achieved ASRs of 0.56 on Llama, 0.35 on Gemma, and 0.66 on Mistral without guards. It remained effective under Forced Refusal Defense, Aligned Model Defense, and Llama Guard 3. In a production GPT-4o fine-tuning experiment using approximately 1,000 examples, NOICE increased ASR from 0.086 to 0.57. The reported supplied data contain no Claude Haiku experiment.

These results imply that refusal training must evaluate the full response trajectory. A response that begins with a refusal but later gives harmful instructions is a safety failure. Output filters must likewise inspect complete responses rather than assigning safety labels from the first sentence or refusal cue.

Refusal composition

Refusal content is heterogeneous. A generic policy denial differs from a truthful statement that the model lacks a modality, information, context, or capability. The human-labeled dataset assembled in (Recum et al., 2024) contains 8,650 input-output pairs, while its multi-annotator subset contains 501 examples. Inter-annotator agreement is moderate, with Krippendorff’s alpha values ranging from 0.497 to 0.608. This indicates that refusal-reason classification is itself ambiguous, particularly for Training Data Limits, Knowledge Cutoff, Missing Context, and Information Hazards.

The dataset includes both human-annotated and synthetic examples. The synthetic collection contains 104,000 examples across 13 categories, with approximately 8,000 examples per category. An expanded variation set contains approximately 7.17 million examples. The classifiers achieve at-least-one agreement of 51.57% with BERT and 78.08% with NV-Embed-V2 plus logistic regression. The reported results support the use of refusal composition analysis for auditing, but do not establish a causal comparison between IFT and RLHF refusal distributions.

3. Data construction and training objectives

Conventional safety fine-tuning

Standard refusal training uses supervised demonstrations in which harmful prompts are paired with safe refusal responses. In autoregressive form, the objective is:

F1F_11

This objective can improve harmful-request refusal, but it may also induce lexical shortcuts and over-refusal. It can associate words such as “kill,” “explode,” or “steal” with refusal even when they occur in technical, metaphorical, or benign contexts.

Decoupled refusal training

DeRTa constructs triples F1F_12 containing a harmful query, a safe response, and a harmful response. A random harmful prefix F1F_13 is inserted into the context, and the model is trained to generate F1F_14:

F1F_15

The complete objective is:

F1F_16

The training pipeline uses 3,000 harmful instructions from BeaverTails, safe responses generated by GPT-3.5-turbo, and harmful responses generated by a maliciously fine-tuned LLaMA3-8B-Instruct model. Harmful responses are edited with GPT-3.5 to improve grammar and lexical diversity and remove safety warnings. Helpfulness training uses 60,000 uncensored Evol-Instruct examples.

Refusal-aware instruction tuning

Refusal-Aware Instruction Tuning, or RAIT, modifies training targets according to whether the initial model answers correctly. Correct examples retain their answers, whereas incorrect examples receive “I don’t know.” “Utilize the Flow before Stepping into the Same River Twice: Certainty Represented Knowledge Flow for Refusal-Aware Instruction Tuning” (Zhu et al., 2024) identifies two sources of over-refusal:

  • Static conflict: similar representations receive contradictory answer and refusal labels.
  • Dynamic conflict: examples initially judged unknown become answerable during training but retain fixed refusal targets.

CRaFT addresses static conflict using response certainty and dynamic conflict using rehearsal training. For each sample, it estimates initial correctness and certainty, then repeats the estimates after rehearsal training:

F1F_17

Initially weak examples whose correctness is increasing under rehearsal are discarded rather than assigned an “I don’t know” target. The remaining examples are ranked by certainty: highly certain answerable examples receive normal answers, while low-certainty persistently unsupported examples receive refusal targets.

In open-ended question answering, certainty is estimated from 10 sampled responses using pairwise cosine similarity between sentence-transformer embeddings. In multiple-choice question answering, certainty is represented by the negative entropy of the option distribution. On TriviaQA, CRaFT improved THS over Cor-RAIT by 3.56 points for LLaMA-2-7B and 2.72 points for LLaMA3-8B. On MMLU, CRaFT improved THS over Cor-RAIT by 1.14 points for LLaMA-2-7B and 13.57 points for LLaMA3-8B.

Refusal tokens

“Refusal Tokens: A Simple Way to Calibrate Refusals in LLMs” (Jain et al., 2024) makes the refusal decision explicit by prepending special tokens to target responses:

F1F_18

Category-specific versions use separate tokens for Humanizing, Indeterminate, Incomplete, Safety, and Unsupported refusals. At inference time, the system can threshold refusal-token probabilities or apply logit bias. A universal refusal token can be emitted when:

F1F_19

A positive logit bias AcatA_{\text{cat}}0 modifies the refusal logit according to:

AcatA_{\text{cat}}1

Category-specific thresholds enable independent control over refusal classes. In the reported experiments, category-token thresholding achieved an F1 score of 0.946, compared with 0.938 under ordinary sampling. Refusal tokens require one fine-tuning stage to establish the token-response association, but subsequent refusal-policy changes do not require further fine-tuning.

The method does not make the token semantically definitive. A refusal token can precede a substantive answer, and a response token can precede a refusal. Contrast examples are therefore necessary to reduce broad refusal generalization.

Information-boundary-aware multimodal training

InBoL combines IDK Instruction Tuning with Confidence-aware Direct Preference Optimization. Its data-generation pipeline begins with VQAv2, Oven, and ScienceQA, estimates confidence from multiple model responses, and generates extrinsically unanswerable questions through mismatched image-question pairs, false premises, and insufficient-information questions.

Using confidence thresholds AcatA_{\text{cat}}2 and AcatA_{\text{cat}}3, examples are divided into known, mixed, and unknown groups. IDK-IT assigns correct answers to known examples and refusal responses to unknown examples while excluding mixed examples. CA-DPO constructs confidence-dependent preferences:

  • high confidence: correct AcatA_{\text{cat}}4 incorrect;
  • low confidence: refusal AcatA_{\text{cat}}5 incorrect;
  • intermediate confidence: both preferences are retained and weighted according to confidence.

On LLaVA1.5-13B, CA-DPO achieved an in-domain trustworthiness score of 32.00, compared with 21.20 for IDK-IT and 17.30 for ordinary SFT. On unanswerable VizWiz questions, IDK-IT and CA-DPO increased refusal rates from 9.60% for the original model to 78.61% and 73.27%, respectively.

Reinforcement finetuning and unanswerable data

“The Hallucination Tax of Reinforcement Finetuning” (Song et al., 20 May 2025) shows that reinforcement finetuning on answerable reasoning problems can substantially reduce refusal behavior. Standard RFT rewards determinate answers but generally provides little or no positive reward for abstention. This encourages hallucinated answers on under-specified problems.

The paper introduces SUM, or Synthetic Unanswerable Math, by modifying answerable competition mathematics problems from DeepScaleR. The five modification types are key-information deletion, ambiguous key information, unrealistic conditions, unrelated objects, and question deletion. A response of \boxed{I don't know.} receives reward 1 on unanswerable questions, while a correct substantive answer receives reward 1 on answerable questions.

Adding 10% SUM increased Qwen2.5-7B-Instruct refusal on UWMP from 0.08 to 0.85 and on SelfAware from 0.09 to 0.99, while GSM8K accuracy decreased from 0.90 to 0.85. The results support the conclusion that unanswerable reasoning examples should be included in RFT, but excessive proportions such as 30% or 50% can reduce answerable-task accuracy.

4. Representation-level refusal mechanisms

Refusal behavior has been modeled as an approximately linear direction in residual-stream activation space. Several methods use difference-in-means vectors between harmful and harmless prompts.

For layer AcatA_{\text{cat}}6, a refusal direction can be estimated as:

AcatA_{\text{cat}}7

The direction can be added to induce refusal or ablated to suppress it:

AcatA_{\text{cat}}8

Refusal Feature Adversarial Training

“Robust LLM safeguarding via refusal feature adversarial training” (Yu et al., 2024) proposes Refusal Feature Adversarial Training, or ReFAT. It interprets many jailbreaks as reducing a refusal-related residual-stream direction, making harmful prompts appear internally more harmless. The refusal feature is estimated from 500 harmful AdvBench instructions and 500 harmless Alpaca instructions.

Refusal Feature Ablation removes the component along the refusal direction and replaces it with the harmless-average projection. The paper reports that GCG, PAIR, AutoDAN, and HumanJailbreaks all produce representation shifts aligned with the negative refusal direction. Restoring the refusal projection sharply reduces attack success rates without meaningful degradation in MMLU or MT-Bench.

ReFAT applies the ablation during training to harmful examples while retaining ordinary utility training on harmless examples. In experiments, it applies RFA over the final 75% of transformer layers. On Llama-3-8B, ReFAT reduced direct RFA attack success from 53.3% for the base model to 10.4%; on Mistral-7B, from 85.6% to 21.1%; and on Gemma-7B, from 92.8% to 18.8%. Its computational overhead was reported as approximately 1.5 forward passes and one backward pass, compared with 11/11 for CAT and 2,566/6 for R2D2.

Single-vector ablation and false refusal

“Surgical, Cheap, and Flexible: Mitigating False Refusal in LLMs via Single Vector Ablation” (Wang et al., 2024) targets a different failure mode: false refusal on benign prompts that resemble harmful prompts. It constructs a raw false-refusal vector from pseudo-harmful prompts and orthogonalizes it against a true-refusal vector. The orthogonalized vector is then ablated from residual activations.

For Llama 2 7B Chat, ablating the raw true-refusal vector increased harmful compliance from 2.3% to 93.0%. Ablating the raw false-refusal vector increased harmful compliance to 46.1%. Ablating the orthogonalized false-refusal vector increased harmful compliance only to 3.1%, while OR-Bench-H compliance rose from 14.8% to 65.6%.

The orthogonalization parameter AcatA_{\text{cat}}9 controls the overlap removed from the false-refusal vector:

v(i,q,r)={1if r is correct, 0if r is a refusal, −1if r is incorrect.v(i,q,r)= \begin{cases} 1 & \text{if } r \text{ is correct},\ 0 & \text{if } r \text{ is a refusal},\ -1 & \text{if } r \text{ is incorrect}. \end{cases}0

Full orthogonalization uses v(i,q,r)={1if r is correct, 0if r is a refusal, −1if r is incorrect.v(i,q,r)= \begin{cases} 1 & \text{if } r \text{ is correct},\ 0 & \text{if } r \text{ is a refusal},\ -1 & \text{if } r \text{ is incorrect}. \end{cases}1. Lower values increase compliance on pseudo-harmful prompts but can also weaken harmful-request refusal. The method is training-free and can be folded into projection weights, leaving inference time and memory unchanged in the reported implementation.

ACTOR

“Just Enough Shifts: Mitigating Over-Refusal in Aligned LLMs with Targeted Representation Fine-Tuning” (Dabas et al., 6 Jul 2025) introduces ACTOR, a response-free method that fine-tunes one selected transformer layer using query-specific activation targets. It identifies the layer with the highest silhouette score separating benign and harmful activations. The selected layers are 13 for Llama-2-7B-Chat, 17 for Gemma-7B-it, and 14 for Llama-2-13B-Chat.

For a query activation v(i,q,r)={1if r is correct, 0if r is a refusal, −1if r is incorrect.v(i,q,r)= \begin{cases} 1 & \text{if } r \text{ is correct},\ 0 & \text{if } r \text{ is a refusal},\ -1 & \text{if } r \text{ is incorrect}. \end{cases}2, ACTOR computes the refusal projection:

v(i,q,r)={1if r is correct, 0if r is a refusal, −1if r is incorrect.v(i,q,r)= \begin{cases} 1 & \text{if } r \text{ is correct},\ 0 & \text{if } r \text{ is a refusal},\ -1 & \text{if } r \text{ is incorrect}. \end{cases}3

Benign and pseudo-harmful queries are shifted away from refusal, while genuinely harmful queries are shifted toward it:

v(i,q,r)={1if r is correct, 0if r is a refusal, −1if r is incorrect.v(i,q,r)= \begin{cases} 1 & \text{if } r \text{ is correct},\ 0 & \text{if } r \text{ is a refusal},\ -1 & \text{if } r \text{ is incorrect}. \end{cases}4

The model minimizes a cosine-alignment loss between current and target activations. Unlike a uniform shift, the intervention magnitude is proportional to each query’s refusal-related component. On Llama-2-7B-Chat, ACTOR increased average compliance from 61.53% to 90.74% while decreasing AdvBench safety score from 99.62% to 99.03%. Training took four minutes on an H100, compared with 15 minutes for the SFT baseline.

Post-training representation drift

“How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence” (Du et al., 3 Apr 2025) compares refusal directions in base, SFT, and instruction-tuned models. It finds that refusal directions differ substantially between base and post-trained models and show limited forward transferability. Native refusal-direction interventions work strongly within the model from which the direction was extracted, but base-model directions often fail to control SFT or instruction-tuned descendants.

For Llama-3.1-8B-Instruct, ablating its native refusal direction reduced harmful-input refusal from 0.98 to 0.01, whereas ablating the base-model direction left refusal at 0.95. Adding the native instruction-tuned direction to harmless inputs raised refusal to 1.00, while adding the base direction raised it only to 0.08. These results indicate that refusal representations can be reorganized during post-training rather than merely thresholded downstream.

“Latent Adversarial Training Improves the Representation of Refusal” (Abbas et al., 26 Apr 2025) examines how latent adversarial training changes refusal geometry. For Llama 2 7B, the first two SVD components explain approximately 74%–75% of harmful-minus-harmless activation variance after LAT, compared with approximately 54.2% for SSFT and 48.6% for embedding-space adversarial training. LAT-derived vectors transfer effectively to other models, but LAT is more vulnerable to self-generated refusal-vector ablation than SSFT and embedding-space adversarial training.

5. Preserving refusal during customization and fine-tuning

Ordinary instruction fine-tuning can damage refusal behavior even when the user data are benign. It can also be deliberately exploited through a small number of harmful examples. “Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint” (Du et al., 8 Sep 2025) attributes this degradation partly to drift in the refusal direction.

ProCon anchors the projection of hidden states onto an initial refusal direction. For each target token and layer, it records an initial projection:

v(i,q,r)={1if r is correct, 0if r is a refusal, −1if r is incorrect.v(i,q,r)= \begin{cases} 1 & \text{if } r \text{ is correct},\ 0 & \text{if } r \text{ is a refusal},\ -1 & \text{if } r \text{ is incorrect}. \end{cases}5

During training, it penalizes deviation from this projection:

v(i,q,r)={1if r is correct, 0if r is a refusal, −1if r is incorrect.v(i,q,r)= \begin{cases} 1 & \text{if } r \text{ is correct},\ 0 & \text{if } r \text{ is a refusal},\ -1 & \text{if } r \text{ is incorrect}. \end{cases}6

The combined loss is:

v(i,q,r)={1if r is correct, 0if r is a refusal, −1if r is incorrect.v(i,q,r)= \begin{cases} 1 & \text{if } r \text{ is correct},\ 0 & \text{if } r \text{ is a refusal},\ -1 & \text{if } r \text{ is incorrect}. \end{cases}7

The enhanced method, v(i,q,r)={1if r is correct, 0if r is a refusal, −1if r is incorrect.v(i,q,r)= \begin{cases} 1 & \text{if } r \text{ is correct},\ 0 & \text{if } r \text{ is a refusal},\ -1 & \text{if } r \text{ is incorrect}. \end{cases}8, applies a strong constraint early in training and then relaxes it. It also adds 1,000 safety-oriented samples to broaden the activation distribution. On LLaMA2-7B, benign IFT increased average ASR from 6.12% to 61.18%, whereas v(i,q,r)={1if r is correct, 0if r is a refusal, −1if r is incorrect.v(i,q,r)= \begin{cases} 1 & \text{if } r \text{ is correct},\ 0 & \text{if } r \text{ is a refusal},\ -1 & \text{if } r \text{ is incorrect}. \end{cases}9 reduced it to 12.88% while increasing task performance from 41.60% to 67.00%. On LLaMA3-8B, it reduced ASR from 71.30% to 9.15%; on Qwen2-7B, from 69.09% to 17.58%.

ReFT addresses a related problem in Finetuning-as-a-Service by filtering harmful user data before supervised fine-tuning and distilling alignment behavior into the student. “Refusal-Feature-guided Teacher for Safe Finetuning via Data Filtering and Alignment Distillation” (Ham et al., 9 Jun 2025) trains a frozen teacher whose last-input-token representation is separated according to a refusal feature. Harmful examples are filtered when their cosine similarity with the refusal feature exceeds Strust=2⋅Acc+RefR−1.S_{\text{trust}}=2\cdot\text{Acc}+\text{RefR}-1.0.

The student objective combines user-task loss on retained examples with KL-divergence distillation from the teacher:

Strust=2⋅Acc+RefR−1.S_{\text{trust}}=2\cdot\text{Acc}+\text{RefR}-1.1

With Llama3-8B and a 50% harmful-data ratio, ordinary SFT reached harmful score 71.3, whereas ReFT reached 0.9. Across Llama3-8B, Gemma2-9B, and Qwen2-7B, ReFT obtained average harmful score 0.8 and finetuning accuracy 60.8. The principal limitation is false-positive filtering: on AlpacaEval, harmful accuracy was 99.9% but harmless accuracy was 77.04%.

6. Evaluation, calibration, and unresolved limitations

Dynamic and selective evaluation

Static direct-prompt benchmarks are insufficient for evaluating refusal training. The past-tense reformulation study “Does Refusal Training in LLMs Generalize to the Past Tense?” (Andriushchenko et al., 2024) shows that a simple linguistic transformation can expose large generalization gaps. Under a GPT-4 judge, GPT-4o’s ASR increased from 1% for direct requests to 88% after 20 past-tense reformulation attempts. With one attempt, the reported ASR was approximately 57%. Future-tense reformulations were less effective, producing 61% ASR for GPT-4o.

The results establish neither that past-tense framing is universally benign nor that data scarcity is the sole causal mechanism. The proposed explanations—historical framing being treated as descriptive and past-tense examples being underrepresented in alignment data—remain hypotheses. Explicitly adding past-tense refusal examples to GPT-3.5 Turbo fine-tuning reduced GPT-4-judged past-tense ASR from 74% to 0% at a 5% refusal-data mixture, but over-refusal increased from 3% to 22%.

RefusalBench (Muhamed et al., 12 Oct 2025) addresses benchmark contamination and narrow test distributions through generative evaluation. It creates 176 perturbation strategies across ambiguity, contradiction, missing information, false premise, granularity mismatch, and epistemic mismatch at low, medium, and high intensity. RefusalBench-NQ contains 1,600 perturbations, while RefusalBench-GaRAGe contains 1,506 multi-document examples.

The results show that selective refusal remains difficult. On NQ, Claude-4-Sonnet achieved the best refusal accuracy at 73.0% and the best CRS at 65.3%. On GaRAGe, DeepSeek-R1 achieved the best refusal accuracy at 47.4%. GPT-4o correctly refused 88.4% of unanswerable NQ cases but had only 54.1% category accuracy and a 62.8% false-refusal rate. Models frequently misclassified ambiguity and granularity mismatch as missing information.

Refusal judges and metrics

Refusal evaluation commonly relies on keyword matching or LLM judges, both of which introduce limitations. “Refusing Safe Prompts for Multi-modal LLMs” (Shao et al., 2024) uses GPT-4 as a refusal judge and detects cues such as “sorry,” “I cannot help,” and “unfortunately.” The main metric is:

Strust=2⋅Acc+RefR−1.S_{\text{trust}}=2\cdot\text{Acc}+\text{RefR}-1.2

This approach can mistake politeness, partial refusal, uncertainty, and evasiveness for genuine safety refusal. Other work uses manually judged ASR, Llama Guard, StrongReject, LlamaGuard-3-8B, Beaver-Dam-7B, or exact refusal strings such as \boxed{I don't know.}. Evaluation should distinguish full refusal, partial refusal, safe redirection, unsupported answer, refusal followed by harmful compliance, and correct answers containing polite caveats.

Safety-guard models introduce an additional failure mode. “When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models” (Feng et al., 4 Aug 2026) audits WildGuardMix and GR-Train and finds that refusal expressions co-occur almost exclusively with unharmful labels. In WildGuardMix, all 10,651 refusal responses associated with harmful prompts are labeled unharmful, with no refusal-plus-harmful examples. In GR-Train, only 114 of 43,074 examples are refusal-plus-harmful.

This imbalance permits a refusal-cue shortcut: inserting a refusal cue into a harmful response can flip the guard’s harmfulness prediction to unharmful. On WildGuardTest, WG-7B exhibited a 37.72% head-position failure rate under at least one of three refusal cues, compared with a 0.71% control rate. The shortcut persisted across head, middle, and tail positions and affected LlamaGuard3 and Qwen3Guard. Sparse complementary masking reduced response-initial detection failures by approximately 79% while preserving standard detection performance.

Principal limitations

Refusal training methods remain constrained by several recurring limitations:

  • Distribution shift: direct harmful prompts, past-tense reformulations, code completion, ciphers, role-play, multi-document contexts, and long-horizon continuations can produce substantially different behavior.
  • Over-refusal: increasing refusal sensitivity can reduce answerable-task accuracy and helpfulness.
  • Synthetic-data artifacts: generated harmful responses, unanswerable questions, and refusal taxonomies may not represent natural user behavior.
  • Metric dependence: keyword matching, exact refusal strings, rule-based judges, and LLM evaluators can disagree about whether a response is genuinely safe or correctly refused.
  • Representation dependence: refusal directions vary by model, layer, token position, architecture, training stage, language, and contrastive dataset.
  • Incomplete causal understanding: linear directions and SVD explain observed activation structure but do not establish that all refusal behavior is one-dimensional or that a single direction captures every safety mechanism.
  • Adaptive attacks: defenses evaluated against fixed attack families may fail against attackers that optimize directly against the defense.
  • Ambiguous normative boundaries: benchmark labels can classify cheating, privacy-sensitive inference, sensitive attributes, or other borderline behaviors inconsistently.
  • Capability trade-offs: strong latent constraints, adversarial training, diffusion purification, and excessive refusal mixtures can impair accuracy, latency, memory, or compute efficiency.
  • Insufficient category coverage: cannot-related refusals, granularity mismatch, ambiguity, missing context, and unknown information are often underrepresented relative to common safety refusals.

The accumulated evidence supports a broad characterization of refusal training as a problem of selective, persistent, and calibrated control. Effective systems must learn when to answer, when to refuse, why refusal is warranted, how to interrupt unsafe generation after it begins, and how to preserve these behaviors during instruction tuning and customization. No single intervention currently addresses all dimensions. Data-centric methods such as DeRTa, CRaFT, InBoL, SUM, and refusal tokens target training distribution and decision calibration; representation-level methods such as ReFAT, single-vector ablation, ACTOR, ReFT, LAT, and ProCon target internal refusal geometry; dynamic benchmarks such as RefusalBench test whether those behaviors generalize beyond fixed examples. Together, these approaches define refusal training as a joint problem of safety alignment, epistemic calibration, representation stability, and behavioral reliability.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (18)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Refusal Training.