Secure Tug-of-War (SecTOW)
- Secure Tug-of-War (SecTOW) is an iterative defense-attack framework designed to enhance MLLM security by countering jailbreaking inputs.
- It employs Group Relative Policy Optimization in alternate attacker and defender training rounds to automatically expand jailbreak data and refine refusal strategies.
- Robust quality monitoring and rule-based rewards yield drastic ASR reductions and low ORR, ensuring both security and general utility.
Secure Tug-of-War (SecTOW) is an iterative defense-attack training framework for enhancing the security of multimodal LLMs (MLLMs) against jailbreaking inputs—maliciously crafted image-query pairs designed to bypass safety constraints. Unlike traditional guardrail-type solutions, which apply external filters but leave internal vulnerabilities unaddressed, or standard supervised fine-tuning (SFT), which often induces excessive refusals on benign queries, SecTOW employs reinforcement learning (specifically, Group Relative Policy Optimization) to conduct an adversarial, end-to-end retraining of both the defender and the attacker. The result is a white-box security protocol that automatically expands jailbreak data, robustly enhances refusal strategies, and monitors both performance and reward exploitation to maintain strong general utility (Dai et al., 29 Jul 2025).
1. System Architecture and Iterative Process
At its core, SecTOW alternates between training two independently parameterized MLLMs—a defender () and an auxiliary attacker ()—in a multiround cycle.
- Defender (): Processes input tuples (image and query), producing a response . Its objective is to safely refuse or answer malicious inputs while retaining informativeness for benign queries.
- Attacker (): Utilizes a “think-then-attack” prompt to generate or refine jailbreak queries, conditioned on images from a general dataset.
Training advances through iterations, each with two phases:
- Attack Phase: is updated via GRPO using current folds of jailbreak () and general data (), generating raw jailbreaks that are filtered for high quality.
- Defense Phase: 0 retrains via GRPO on a union of newly discovered jailbreaks and benign samples, updating its refusal strategy.
Data partitioning is performed by splitting both the initial, scarce jailbreak dataset 1 and a much larger general set 2 into 3 folds, cycling through each fold per iteration. Cold-start phases (SFT for 4 and initial GRPO for 5) address reward sparsity at the outset (Dai et al., 29 Jul 2025).
2. Reinforcement Learning Formalism and GRPO
Both defender and attacker operate as autoregressive policies 6, generating token sequences over the respective state and action spaces:
- State space 7: 8, where 9 is the sequence of previously generated tokens.
- Action space 0: MLLM vocabulary.
Policies are optimized to maximize expected return:
1
where trajectories 2 accrue cumulative reward 3.
SecTOW implements Group Relative Policy Optimization (GRPO), which eschews a value baseline in favor of intra-group relative advantage:
- For each input, 4 rollouts 5 produce rewards 6; mean 7 and advantage 8 are computed.
- The policy is updated using the surrogate loss:
9
where 0 is a KL regularization coefficient, often tuned dynamically.
GRPO is favored for its simplicity (no learned value function), direct regularization, and variance reduction within large-scale MLLM training (Dai et al., 29 Jul 2025).
3. Reward Design: Defender and Attacker Objectives
Distinct reward functions are utilized for each role, built on simple, deterministic criteria rather than model-based or generative labeling.
- Defender Reward: For each sample 1, assign binary label 2 (“reject required” or “pass through”). Rule-based detection 3 establishes if the model’s response matches a refusal phrase. Reward function:
4
This yields stable, low-noise feedback.
- Attacker Reward: For each (image, query) 5, a composite criterion is used. First, pass the query through the defender and judge if the response is unsafe (via an external safety evaluator). Second, enforce template adherence in the generated query:
6
where both scores are binary indicators.
These rule-based mechanisms reduce the need for complex generative annotations and enable robust scaling to synthetic samples (Dai et al., 29 Jul 2025).
4. Quality Monitoring to Mitigate Reward Gaming
To avoid degenerate behaviors—such as the attacker generating repetitive, low-diversity queries or the defender excessively rejecting benign inputs (“over-refusal”)—SecTOW incorporates explicit early stopping criteria:
- Attacker Diversity: On a validation set, computes
7
(using FuzzyWuzzy or other string similarity). If 8 drops by 9, attacker retraining is halted.
- Defender Over-Refusal Rate (ORR): On general validation,
0
If ORR exceeds 1, defender retraining is stopped to preserve utility.
This systematic monitoring ensures that neither model exploits the rule-based rewards to the detriment of system balance or general purpose performance (Dai et al., 29 Jul 2025).
5. Performance Evaluation and Ablation Analysis
SecTOW was instantiated with Qwen2-VL-7B for both defender and attacker, using Llama-Guard-3 as the external safety evaluator. Benchmarks include:
- Safety-specific: JailBreakV-28k, FigStep, SafeBench, MM-SafetyBench.
- General: MMMU, MMMU-Pro.
Key metrics:
- Attack Success Rate (ASR): Fraction of jailbreak inputs that elicit unintended outputs.
- Accuracy and Over-Refusal Rate (ORR): On standard multimodal benchmarks.
Post three SecTOW iterations, results are as follows:
| Benchmark | Initial ASR / ACC∕ORR | Post-SecTOW ASR / ACC∕ORR |
|---|---|---|
| JailBreakV-28k | 0.1918 | 0.0061 |
| FigStep | 0.3320 | 0.0000 |
| SafeBench | 0.1404 | 0.0022 |
| MM-SafetyBench | 0.6726 | 0.0298 |
| MMMU | 0.5411∕0.0000 | 0.5422∕0.0078 |
| MMMU-Pro | 0.4116∕0.0000 | 0.4145∕0.0006 |
Compared to SFT, which reduced ASR but increased ORR above 20%, SecTOW achieves lower ASR while keeping ORR beneath 2. Relative to other dense defenses (MMS, MIRage), SecTOW reduces ASR by at least 88% on JailBreakV-28k and neutralizes all FigStep attacks (Dai et al., 29 Jul 2025).
Ablation studies demonstrate the necessity of iteration, defender and attacker monitoring, and cold start phases for optimal robustness and usability. Notably, disabling diversity monitoring or defender ORR checks led to substantial ASR and ORR degradations, respectively.
6. Implementation Notes and Adaptation Strategies
SecTOW’s reproducibility and adaptability to diverse MLLM contexts are underpinned by modular and lightweight design choices:
- Model and Evaluator Choices: Qwen2-VL-7B and Llama-Guard-3; suitable substitutions can be made for other architectures (e.g., Flamingo, LLaVA).
- Iteration Count: 3; attacker allocation: 80% training, 20% validation for diversity.
- Cold Start: Defender SFT on full seed data; attacker early GRPO warm-up.
- GRPO Settings: Learning rate 4, group size 5, KL weight 6.
- Jailbreak Filtering: Retain generated (image, query) pairs if at least 7 out of 8 defender samples are successful attacks.
- Integration: Adaptation requires (a) swapping in the target architecture, (b) preparing minimal jailbreak and general sets, (c) implementing GRPO-based training, and (d) leveraging rule-based reward and monitoring routines as specified.
The minimal reliance on learned rewards and annotations, coupled with explicit cycle monitoring, makes SecTOW broadly extensible for robust, scalable MLLM security enhancement (Dai et al., 29 Jul 2025).