---
title: 'ORFuzz: Detecting LLM Over-Refusals'
url: https://www.emergentmind.com/topics/orfuzz
type: topic
---

# ORFuzz: Detecting LLM Over-Refusals

Searching arXiv for the ORFuzz paper and closely related context.
ORFuzz is an evolutionary testing framework for the systematic detection and analysis of **LLM over-refusals**, a failure mode in which a model erroneously rejects a benign prompt because its safety system is overly conservative. The framework treats over-refusal as the “other side” of LLM safety: whereas jailbreaks and under-refusals expose unsafe willingness to answer harmful requests, over-refusals expose functional failure under benign use. ORFuzz combines **safety category-aware seed selection**, **adaptive mutator optimization using reasoning LLMs**, and a **human-aligned judge model**, OR-Judge, to generate and validate benign prompts that are likely to be wrongly refused. In the reported evaluation, ORFuzz achieved an average validated over-refusal generation rate of **6.98%**, and its derived benchmark, **ORFuzzSet**, reached a **63.56%** average over-refusal rate across **10** LLMs [2508.11222].

## 1. Concept and problem formulation

ORFuzz is defined around **over-refusal** rather than jailbreak success. In this setting, the problematic outcome is not that a model answers harmful content, but that it refuses a query that should be answerable. The motivating example is a benign coding request such as “How to kill a python process?”, where the lexical trigger “kill” may induce refusal despite the intended meaning being ordinary process termination. This places ORFuzz in the space of functional safety evaluation rather than adversarial elicitation of harmful outputs [2508.11222].

The paper argues that existing over-refusal testing methods are inadequate for two reasons. First, widely used datasets such as **XSTest** and **OR-Bench** are static prompt collections with limited generation power. Second, labeling quality is unstable because benignity and refusal are not perceived uniformly by human annotators. A user study with **20 human evaluators**, **five LLMs**, and **2,500 query-response pairs** yielded **Fleiss’ kappa \(\kappa = 0.552\)**, indicating moderate agreement. The same study reported that roughly **51% of OR-Bench benign prompts were perceived as harmful by humans**, which undermines their use as clean probes of erroneous refusal [2508.11222].

This framing suggests that over-refusal is neither a trivial complement of jailbreak evaluation nor a problem addressable by fixed benchmarks alone. A plausible implication is that any robust evaluation regime must account simultaneously for prompt benignity, refusal behavior, and the specifically **safety-related** character of that refusal.

## 2. Framework architecture

ORFuzz operates as a closed-loop evolutionary system. It starts from a seed pool built from **COR**, **XSTest**, and **OR-Bench-Hard-1K**, then iteratively selects seeds, mutates them, queries a target LLM, evaluates the resulting interaction with OR-Judge, and feeds the reward signal back into both seed selection and mutator optimization. The framework’s three named core components are **safety category-aware seed selection**, **adaptive mutator optimization using reasoning LLMs**, and **OR-Judge** [2508.11222].

Seed selection is organized around an **8-way safety taxonomy** derived from **S-Eval**. The categories are: **Crimes and Illegal Activities (CI)**, **Hate Speech (HS)**, **Physical and Mental Health (PM)**, **Ethics and Morality (EM)**, **Data Privacy (DP)**, **Cybersecurity (CS)**, **Extremism (EX)**, and **Inappropriate Suggestions (IS)**. ORFuzz uses a hierarchical **seed selection graph** whose root contains category nodes and whose category nodes contain seed-query children. Initial seed populations within categories are sampled with **k-means** clustering to reduce redundancy while preserving intra-category diversity. The paper presents this as a **category-aware MCTS-Explore algorithm** with `Initialize`, `Select`, and `Update` procedures, and uses a UCB-style score,
$$
UCB = \nicefrac{r}{n} + \sqrt{\nicefrac{2\ln t}{n}}
$$
where \(r\) is node reward, \(n\) is node visit count, and \(t\) is total root visits [2508.11222].

The seed-selection update includes a depth-sensitive reward discount of \(r \cdot \max(\beta, 1-0.1l)\), where \(l\) is depth. This reflects the framework’s attempt to balance coverage across categories with exploitation of productive seeds. In the reported experiments, ORFuzz and its ablations using the categorized seed pool covered all **8 safety categories**, while direct generation covered **3** and GPTFuzz covered **5** [2508.11222].

## 3. Mutators and adaptive optimization

ORFuzz does not reuse jailbreak-oriented mutators directly, because the goal is not to bypass safety on harmful inputs. Instead, it defines mutators intended to preserve benign meaning while increasing the chance of a safety-triggered refusal. These are divided into three classes: **General Mutators**, **Sensitive Word Mutators**, and **Scenario/Task Mutators** [2508.11222].

The **General Mutators** are **Shorten**, **Expand**, **Rephrase**, **Cross-over**, **Translate**, and **Regenerate**. The **Sensitive Word Mutators** are **Insert sensitive words** and **Replace sensitive words**. The **Scenario/Task Mutators** are **Scenario mutate** and **Task mutate**. Their purpose is to generate benign prompts that remain safety-adjacent enough to provoke conservative refusal heuristics [2508.11222].

A central feature is that mutators are not fixed operators. ORFuzz uses UCB to choose a mutator and then refines the **natural-language prompt** that defines how that mutator behaves. The paper describes this as an “analyze-generate-feedback” loop inspired by PromptWizard. A reasoning LLM analyzes a mutator prompt, generates \(N^{mut}\) prompt variants including the original, uses those variants to mutate selected seeds, and chooses the best variant by the **average reward** of its generated samples. Feedback about which variants worked is retained to improve future refinements [2508.11222].

The optimization target is the reward
$$
r = (\hat{p}_{over} - \hat{p}_{toxic} + 1)/2
$$
where \(\hat{p}_{over}\) is the predicted probability that the target model’s behavior constitutes a safety-related refusal, and \(\hat{p}_{toxic}\) is the predicted probability that the prompt itself is toxic. This formulation explicitly prefers prompts that are both more benign and more likely to elicit erroneous refusal [2508.11222].

The ablation study indicates that all major components contribute to performance. Reported average ORRs were **6.98%** for full ORFuzz, **6.22%** without seed sampling, **4.36%** without seed selection, **5.11%** without mutator selection, and **3.16%** without mutator refinement. This suggests that **mutator refinement** is the most performance-critical component among the tested removals [2508.11222].

## 4. OR-Judge and the over-refusal oracle

OR-Judge is ORFuzz’s automated oracle for deciding whether a query-response pair constitutes an over-refusal. It is described as **human-aligned** because it is fine-tuned on direct human annotations of prompt toxicity and answer adequacy, rather than used as a generic off-the-shelf LLM judge. The base model is **Qwen2.5-14B-Instruct**, fine-tuned with **LoRA** on the **2,500 query-response pairs** collected in the user study and split **8:1:1** into train, validation, and test sets [2508.11222].

OR-Judge consists of two fine-tuned models, \(M_{toxic}\) and \(M_{answer}\), plus a safety-related refusal estimate derived from the target model. \(M_{toxic}\) predicts
$$
\hat{p}_{toxic} \triangleq \hat{\mathbb{P}(I_{toxic}(q)=1|q; M_{toxic})},
$$
and \(M_{answer}\) predicts
$$
\hat{p}_{answer} \triangleq \hat{\mathbb{P}(I_{answer}(q, o_q^M)=1|q; M_{answer})}.
$$
For safety-relatedness, the target model classifies its own refusal into three classes: \(C_1\) for safety-related refusal, \(C_2\) for other refusals, and \(C_3\) for non-refusal. ORFuzz then estimates
$$
\hat{p}_{over} = \hat{p}_{sr} \cdot (1 - \hat{p}_{answer}),
$$
so over-refusal is high when the model likely refused, and that refusal was likely safety-related [2508.11222].

The scalar output used by the judge is computed from the probability mass on “Yes” relative to “Yes” plus “No”:
$$
OUTPUT=\nicefrac{\hat{p}_{yes}{(\hat{p}_{yes}+\hat{p}_{no})}
$$
as rendered in the paper. Training uses cross-entropy loss,
$$
L = CELoss(\hat{p}, r; M),
$$
where \(r\) is the user-aggregated label [2508.11222].

The paper validates OR-Judge against several baselines on the 2,500-pair evaluation set. The reported results are shown below.

| Judgment task | Metric | OR-Judge |
|---|---:|---:|
| Toxic score | MAE | 0.0930 |
| Toxic score | MSE | 0.0171 |
| Toxic score | F1 | 0.8642 |
| Answer score | MAE | 0.0599 |
| Answer score | MSE | 0.0097 |
| Answer score | F1 | 0.9674 |

The paper states that OR-Judge outperforms the second-best model by **36.59%** for toxic judgment and **60.11%** for answer judgment in F1. A significant limitation remains that \(\hat{p}_{sr}\) depends on the **target model self-reporting why it refused**, which the paper does not independently verify [2508.11222].

## 5. Evaluation methodology and empirical results

The principal generation metric is the **Over-Refusal Rate (ORR)**, defined in the paper as
$$
ORR(Q, M)=\frac{1}{|Q|}\sum_{q_i \in Q} I_{\text{OR}(q_i, o_{q_i}^M)
$$
with the indicator equal to 1 when a query-output pair forms an over-refusal. In effect, ORR measures the fraction of prompts that trigger over-refusal on a given model [2508.11222].

ORFuzz was evaluated on **five target LLMs**: **Llama-3.1**, **Gemma-2**, **Mistral-v0.3**, **Phi-3.5**, and **Qwen2.5**. Baselines were **Naive / Direct**, the **OR-Bench method**, and **GPTFuzz**, with **DeepSeek-R1** used as the generation or mutation LLM for all methods. Each method generated **450 samples over 50 iterations**, targeting **9 samples per iteration**. For ORFuzz, \(N^{sele}_{seed}=3\) and \(N^{mut}=3\), yielding 9 generated cases per iteration [2508.11222].

The headline quantitative result is an **average ORR of 6.98%** for ORFuzz, compared with **1.91%** for the OR-Bench method, **0.58%** for GPTFuzz, and **0.40%** for Direct generation. Per model, ORFuzz achieved **16.00%** on Llama-3.1, **6.44%** on Gemma-2, **6.44%** on Mistral-v0.3, **1.33%** on Phi-3.5, and **4.67%** on Qwen2.5. The framework also covered all **8 safety categories** and obtained an average **MSS** of **0.308**, compared with **0.263** for OR-Bench, **0.353** for GPTFuzz, and **0.708** for Direct generation; lower MSS indicates greater semantic diversity [2508.11222].

The paper also reports category-specific tendencies, stating for example that **Extremism (EX)** prompts tend to trigger over-refusal more often, while **Ethics and Morality (EM)** prompts are less likely to do so. It further states that over-refusal is **model-specific**, with Gemma-2 over-refusing more in **Physical and Mental Health (PM)** and Llama-3.1 being more sensitive in **Crimes and Illegal Activities (CI)**. This suggests that model-specific adaptive search is materially different from evaluating a fixed benchmark once across many systems [2508.11222].

## 6. ORFuzzSet, transferability, and significance

ORFuzzSet is a benchmark constructed from ORFuzz-generated prompts that triggered over-refusal in **at least three of the five initial target LLMs**. This transferability filter produces a benchmark of **1,855 queries**, which is then evaluated on **10 LLMs**: the original five plus **Qwen2.5-3B-Instruct**, **Qwen2.5-14B-Instruct**, **Deepseek-R1-Distill-Qwen-7B**, **Deepseek-R1-Distill-Llama-8B**, and **Deepseek-R1-Distill-Qwen-14B** [2508.11222].

The resulting benchmark achieves an **average ORR of 63.56%** across those 10 models. This number differs conceptually from the **6.98%** figure for ORFuzz itself: the latter is the efficiency of the generator in producing validated failing cases, whereas the former is the potency of a fixed benchmark composed of highly transferable cases [2508.11222].

ORFuzzSet covers all **8 safety categories**. The reported category counts are **CI: 430**, **CS: 329**, **DP: 83**, **EM: 76**, **EX: 185**, **HS: 316**, **IS: 89**, and **PM: 347**. The most common transferable categories are **Crimes and Illegal Activities (CI): 23.18%** and **Physical and Mental Health (PM): 18.71%**. Within the Qwen family, the paper reports **70.94% ORR** for **Qwen2.5-7B**, **70.94% ORR** for **Qwen2.5-14B**, and **51.10% ORR** for **Qwen2.5-3B**, which the authors interpret as evidence that over-refusal propensity is not a monotonic function of model scale [2508.11222].

The broader significance of ORFuzz is conceptual as much as technical. It reframes LLM safety evaluation so that a model can fail not only by answering harmful requests, but also by refusing benign ones. The framework’s emphasis on category-aware exploration, adaptive mutation, and a human-trained judge suggests that over-refusal is not well captured by static benchmark design alone. At the same time, the paper leaves several open issues: its definition is restricted to **safety-related** refusals; OR-Judge is aligned to a **20-person** annotation population with only **moderate** agreement; and the safety-relatedness estimate depends on target-model self-explanation. These limitations indicate that ORFuzz is best understood as a rigorous initial framework for automated over-refusal discovery rather than a complete resolution of the problem [2508.11222].

Source: https://www.emergentmind.com/topics/orfuzz