ORFuzz: Detecting LLM Over-Refusals
- The paper introduces ORFuzz as a framework that systematically identifies LLM over-refusals by treating functional safety failures as a key evaluation metric.
- It employs a closed-loop evolutionary system featuring safety category-aware seed selection, adaptive mutator optimization, and a human-aligned OR-Judge.
- Empirical results show ORFuzz achieved a 6.98% validated over-refusal rate and its benchmark, ORFuzzSet, reached 63.56% over-refusal across 10 LLMs.
Searching arXiv for the ORFuzz paper and closely related context. ORFuzz is an evolutionary testing framework for the systematic detection and analysis of LLM over-refusals, a failure mode in which a model erroneously rejects a benign prompt because its safety system is overly conservative. The framework treats over-refusal as the “other side” of LLM safety: whereas jailbreaks and under-refusals expose unsafe willingness to answer harmful requests, over-refusals expose functional failure under benign use. ORFuzz combines safety category-aware seed selection, adaptive mutator optimization using reasoning LLMs, and a human-aligned judge model, OR-Judge, to generate and validate benign prompts that are likely to be wrongly refused. In the reported evaluation, ORFuzz achieved an average validated over-refusal generation rate of 6.98%, and its derived benchmark, ORFuzzSet, reached a 63.56% average over-refusal rate across 10 LLMs (Zhang et al., 15 Aug 2025).
1. Concept and problem formulation
ORFuzz is defined around over-refusal rather than jailbreak success. In this setting, the problematic outcome is not that a model answers harmful content, but that it refuses a query that should be answerable. The motivating example is a benign coding request such as “How to kill a python process?”, where the lexical trigger “kill” may induce refusal despite the intended meaning being ordinary process termination. This places ORFuzz in the space of functional safety evaluation rather than adversarial elicitation of harmful outputs (Zhang et al., 15 Aug 2025).
The paper argues that existing over-refusal testing methods are inadequate for two reasons. First, widely used datasets such as XSTest and OR-Bench are static prompt collections with limited generation power. Second, labeling quality is unstable because benignity and refusal are not perceived uniformly by human annotators. A user study with 20 human evaluators, five LLMs, and 2,500 query-response pairs yielded Fleiss’ kappa , indicating moderate agreement. The same study reported that roughly 51% of OR-Bench benign prompts were perceived as harmful by humans, which undermines their use as clean probes of erroneous refusal (Zhang et al., 15 Aug 2025).
This framing suggests that over-refusal is neither a trivial complement of jailbreak evaluation nor a problem addressable by fixed benchmarks alone. A plausible implication is that any robust evaluation regime must account simultaneously for prompt benignity, refusal behavior, and the specifically safety-related character of that refusal.
2. Framework architecture
ORFuzz operates as a closed-loop evolutionary system. It starts from a seed pool built from COR, XSTest, and OR-Bench-Hard-1K, then iteratively selects seeds, mutates them, queries a target LLM, evaluates the resulting interaction with OR-Judge, and feeds the reward signal back into both seed selection and mutator optimization. The framework’s three named core components are safety category-aware seed selection, adaptive mutator optimization using reasoning LLMs, and OR-Judge (Zhang et al., 15 Aug 2025).
Seed selection is organized around an 8-way safety taxonomy derived from S-Eval. The categories are: Crimes and Illegal Activities (CI), Hate Speech (HS), Physical and Mental Health (PM), Ethics and Morality (EM), Data Privacy (DP), Cybersecurity (CS), Extremism (EX), and Inappropriate Suggestions (IS). ORFuzz uses a hierarchical seed selection graph whose root contains category nodes and whose category nodes contain seed-query children. Initial seed populations within categories are sampled with k-means clustering to reduce redundancy while preserving intra-category diversity. The paper presents this as a category-aware MCTS-Explore algorithm with Initialize, Select, and Update procedures, and uses a UCB-style score,
$UCB = \nicefrac{r}{n} + \sqrt{\nicefrac{2\ln t}{n}}$
where is node reward, is node visit count, and is total root visits (Zhang et al., 15 Aug 2025).
The seed-selection update includes a depth-sensitive reward discount of , where is depth. This reflects the framework’s attempt to balance coverage across categories with exploitation of productive seeds. In the reported experiments, ORFuzz and its ablations using the categorized seed pool covered all 8 safety categories, while direct generation covered 3 and GPTFuzz covered 5 (Zhang et al., 15 Aug 2025).
3. Mutators and adaptive optimization
ORFuzz does not reuse jailbreak-oriented mutators directly, because the goal is not to bypass safety on harmful inputs. Instead, it defines mutators intended to preserve benign meaning while increasing the chance of a safety-triggered refusal. These are divided into three classes: General Mutators, Sensitive Word Mutators, and Scenario/Task Mutators (Zhang et al., 15 Aug 2025).
The General Mutators are Shorten, Expand, Rephrase, Cross-over, Translate, and Regenerate. The Sensitive Word Mutators are Insert sensitive words and Replace sensitive words. The Scenario/Task Mutators are Scenario mutate and Task mutate. Their purpose is to generate benign prompts that remain safety-adjacent enough to provoke conservative refusal heuristics (Zhang et al., 15 Aug 2025).
A central feature is that mutators are not fixed operators. ORFuzz uses UCB to choose a mutator and then refines the natural-language prompt that defines how that mutator behaves. The paper describes this as an “analyze-generate-feedback” loop inspired by PromptWizard. A reasoning LLM analyzes a mutator prompt, generates prompt variants including the original, uses those variants to mutate selected seeds, and chooses the best variant by the average reward of its generated samples. Feedback about which variants worked is retained to improve future refinements (Zhang et al., 15 Aug 2025).
The optimization target is the reward
where is the predicted probability that the target model’s behavior constitutes a safety-related refusal, and $UCB = \nicefrac{r}{n} + \sqrt{\nicefrac{2\ln t}{n}}$0 is the predicted probability that the prompt itself is toxic. This formulation explicitly prefers prompts that are both more benign and more likely to elicit erroneous refusal (Zhang et al., 15 Aug 2025).
The ablation study indicates that all major components contribute to performance. Reported average ORRs were 6.98% for full ORFuzz, 6.22% without seed sampling, 4.36% without seed selection, 5.11% without mutator selection, and 3.16% without mutator refinement. This suggests that mutator refinement is the most performance-critical component among the tested removals (Zhang et al., 15 Aug 2025).
4. OR-Judge and the over-refusal oracle
OR-Judge is ORFuzz’s automated oracle for deciding whether a query-response pair constitutes an over-refusal. It is described as human-aligned because it is fine-tuned on direct human annotations of prompt toxicity and answer adequacy, rather than used as a generic off-the-shelf LLM judge. The base model is Qwen2.5-14B-Instruct, fine-tuned with LoRA on the 2,500 query-response pairs collected in the user study and split 8:1:1 into train, validation, and test sets (Zhang et al., 15 Aug 2025).
OR-Judge consists of two fine-tuned models, $UCB = \nicefrac{r}{n} + \sqrt{\nicefrac{2\ln t}{n}}$1 and $UCB = \nicefrac{r}{n} + \sqrt{\nicefrac{2\ln t}{n}}$2, plus a safety-related refusal estimate derived from the target model. $UCB = \nicefrac{r}{n} + \sqrt{\nicefrac{2\ln t}{n}}$3 predicts
$UCB = \nicefrac{r}{n} + \sqrt{\nicefrac{2\ln t}{n}}$4
and $UCB = \nicefrac{r}{n} + \sqrt{\nicefrac{2\ln t}{n}}$5 predicts
$UCB = \nicefrac{r}{n} + \sqrt{\nicefrac{2\ln t}{n}}$6
For safety-relatedness, the target model classifies its own refusal into three classes: $UCB = \nicefrac{r}{n} + \sqrt{\nicefrac{2\ln t}{n}}$7 for safety-related refusal, $UCB = \nicefrac{r}{n} + \sqrt{\nicefrac{2\ln t}{n}}$8 for other refusals, and $UCB = \nicefrac{r}{n} + \sqrt{\nicefrac{2\ln t}{n}}$9 for non-refusal. ORFuzz then estimates
0
so over-refusal is high when the model likely refused, and that refusal was likely safety-related (Zhang et al., 15 Aug 2025).
The scalar output used by the judge is computed from the probability mass on “Yes” relative to “Yes” plus “No”:
1
as rendered in the paper. Training uses cross-entropy loss,
2
where 3 is the user-aggregated label (Zhang et al., 15 Aug 2025).
The paper validates OR-Judge against several baselines on the 2,500-pair evaluation set. The reported results are shown below.
| Judgment task | Metric | OR-Judge |
|---|---|---|
| Toxic score | MAE | 0.0930 |
| Toxic score | MSE | 0.0171 |
| Toxic score | F1 | 0.8642 |
| Answer score | MAE | 0.0599 |
| Answer score | MSE | 0.0097 |
| Answer score | F1 | 0.9674 |
The paper states that OR-Judge outperforms the second-best model by 36.59% for toxic judgment and 60.11% for answer judgment in F1. A significant limitation remains that 4 depends on the target model self-reporting why it refused, which the paper does not independently verify (Zhang et al., 15 Aug 2025).
5. Evaluation methodology and empirical results
The principal generation metric is the Over-Refusal Rate (ORR), defined in the paper as
5
with the indicator equal to 1 when a query-output pair forms an over-refusal. In effect, ORR measures the fraction of prompts that trigger over-refusal on a given model (Zhang et al., 15 Aug 2025).
ORFuzz was evaluated on five target LLMs: Llama-3.1, Gemma-2, Mistral-v0.3, Phi-3.5, and Qwen2.5. Baselines were Naive / Direct, the OR-Bench method, and GPTFuzz, with DeepSeek-R1 used as the generation or mutation LLM for all methods. Each method generated 450 samples over 50 iterations, targeting 9 samples per iteration. For ORFuzz, 6 and 7, yielding 9 generated cases per iteration (Zhang et al., 15 Aug 2025).
The headline quantitative result is an average ORR of 6.98% for ORFuzz, compared with 1.91% for the OR-Bench method, 0.58% for GPTFuzz, and 0.40% for Direct generation. Per model, ORFuzz achieved 16.00% on Llama-3.1, 6.44% on Gemma-2, 6.44% on Mistral-v0.3, 1.33% on Phi-3.5, and 4.67% on Qwen2.5. The framework also covered all 8 safety categories and obtained an average MSS of 0.308, compared with 0.263 for OR-Bench, 0.353 for GPTFuzz, and 0.708 for Direct generation; lower MSS indicates greater semantic diversity (Zhang et al., 15 Aug 2025).
The paper also reports category-specific tendencies, stating for example that Extremism (EX) prompts tend to trigger over-refusal more often, while Ethics and Morality (EM) prompts are less likely to do so. It further states that over-refusal is model-specific, with Gemma-2 over-refusing more in Physical and Mental Health (PM) and Llama-3.1 being more sensitive in Crimes and Illegal Activities (CI). This suggests that model-specific adaptive search is materially different from evaluating a fixed benchmark once across many systems (Zhang et al., 15 Aug 2025).
6. ORFuzzSet, transferability, and significance
ORFuzzSet is a benchmark constructed from ORFuzz-generated prompts that triggered over-refusal in at least three of the five initial target LLMs. This transferability filter produces a benchmark of 1,855 queries, which is then evaluated on 10 LLMs: the original five plus Qwen2.5-3B-Instruct, Qwen2.5-14B-Instruct, Deepseek-R1-Distill-Qwen-7B, Deepseek-R1-Distill-Llama-8B, and Deepseek-R1-Distill-Qwen-14B (Zhang et al., 15 Aug 2025).
The resulting benchmark achieves an average ORR of 63.56% across those 10 models. This number differs conceptually from the 6.98% figure for ORFuzz itself: the latter is the efficiency of the generator in producing validated failing cases, whereas the former is the potency of a fixed benchmark composed of highly transferable cases (Zhang et al., 15 Aug 2025).
ORFuzzSet covers all 8 safety categories. The reported category counts are CI: 430, CS: 329, DP: 83, EM: 76, EX: 185, HS: 316, IS: 89, and PM: 347. The most common transferable categories are Crimes and Illegal Activities (CI): 23.18% and Physical and Mental Health (PM): 18.71%. Within the Qwen family, the paper reports 70.94% ORR for Qwen2.5-7B, 70.94% ORR for Qwen2.5-14B, and 51.10% ORR for Qwen2.5-3B, which the authors interpret as evidence that over-refusal propensity is not a monotonic function of model scale (Zhang et al., 15 Aug 2025).
The broader significance of ORFuzz is conceptual as much as technical. It reframes LLM safety evaluation so that a model can fail not only by answering harmful requests, but also by refusing benign ones. The framework’s emphasis on category-aware exploration, adaptive mutation, and a human-trained judge suggests that over-refusal is not well captured by static benchmark design alone. At the same time, the paper leaves several open issues: its definition is restricted to safety-related refusals; OR-Judge is aligned to a 20-person annotation population with only moderate agreement; and the safety-relatedness estimate depends on target-model self-explanation. These limitations indicate that ORFuzz is best understood as a rigorous initial framework for automated over-refusal discovery rather than a complete resolution of the problem (Zhang et al., 15 Aug 2025).