Explain Transferability of Single-Query-Trained BPJ Attacks
Investigate, both formally and empirically, why adversarial prefixes learned by Boundary Point Jailbreaking (BPJ) when optimized on a single harmful target query (i.e., in the single-attack setting) generalize and transfer to a wide range of unseen queries.
References
A range of open questions remain, including developing defences to BPJ and exploring formally and empirically why BPJ attacks learned on a single attack readily transfer to other queries.
Although our experiments show strong ASR improvements, transferability remains inherently uncertain. When source and target models differ substantially in architecture, training data, alignment procedure, or refusal behavior, source-side diagnostics may fail to predict target-model effectiveness.