Enhancing Jailbreak Attacks with Diversity Guidance (2403.00292v2)

Published 1 Mar 2024 in cs.CL

Abstract: As LLMs(LLMs) become commonplace in practical applications, the security issues of LLMs have attracted societal concerns. Although extensive efforts have been made to safety alignment, LLMs remain vulnerable to jailbreak attacks. We find that redundant computations limit the performance of existing jailbreak attack methods. Therefore, we propose DPP-based Stochastic Trigger Searching (DSTS), a new optimization algorithm for jailbreak attacks. DSTS incorporates diversity guidance through techniques including stochastic gradient search and DPP selection during optimization. Detailed experiments and ablation studies demonstrate the effectiveness of the algorithm. Moreover, we use the proposed algorithm to compute the risk boundaries for different LLMs, providing a new perspective on LLM safety evaluation.

References (53)

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Enhancing Jailbreak Attacks with Diversity Guidance (2403.00292v2)

Summary

Related Papers

Tweets